VLM
适配强。原片 + 叙述即可开训 / 评测;手与检测为可选辅助。
Strong fit. Source + narration is enough to train or eval; hands and detectors stay optional helpers.
TOBI
原生帧率原片、文本时间戳、带偏移的音频、佩戴者视角的人工句子。设备测量与后处理估计分开。这是人佩戴观测,不是机器人动作包。
Native-Hz source, text timestamps, audio with its offset, and wearer-scoped human sentences. Device measurements and later estimates stay apart. Human-worn observation — not a robot action pack.
主推方案Primary solution
第一人称 VLM 主要读四样:原生帧率画面、文本时间戳、带偏移的音频、佩戴者视角的人工句子。E6 / E2 按 device / human / indexes 交付这四样。手轨迹与 IMU 可作条件,不写成机器人动作。世界模型与短程 VLN 仍是旁路;标准 VLA 仍缺末端与接触力。
Egocentric VLMs mainly read four things: native-Hz frames, text timestamps, audio with its offset, and wearer-scoped human sentences. E6 / E2 deliver those four under device / human / indexes. Hand tracks and IMU may condition a model. They are not robot actions. World models and short VLN stay side paths; standard VLA still needs an end-effector and contact force.
适配强。原片 + 叙述即可开训 / 评测;手与检测为可选辅助。
Strong fit. Source + narration is enough to train or eval; hands and detectors stay optional helpers.
可部分用。长双目与多相机适合视频预测;缺可交互环境状态。
Partial. Long stereo and multi-cam suit video prediction; no interactive world state.
可部分用(表征)。E6 有头部位姿;无地图与标准 path–指令对。
Partial (representation). E6 has head pose; no map or standard path–instruction pairs.
当前不适配。无末端、夹爪、接触力;手柄位姿为零行。
Not a fit today. No eef, gripper, or contact force; controller pose is a zero row.
第一人称 VQA、场景描述、手–物理解:视觉读 device 原片,语言读 human,时间读 indexes。工作台六步句可直接核对。
Egocentric VQA, scene captioning, hand–object understanding: vision from device source, language from human, time from indexes. The six workbench sentences are checkable as-is.
对齐时钟 → 按索引取原生帧 → 挂叙述时间段 → 写成文本时间戳并写上音频偏移 → 可选挂手点或隐私掩码。预览 mp4 只给人看。
Align clocks → seek native frames by index → attach narration spans → write a text timestamp and the audio offset → optionally attach hand points or privacy masks. Preview mp4 is for humans only.
不写成 VLA-ready。不把 OpenXR / Vision 手当成 robot action。不把 SGBM 距离当成深度相机真值。不把论文 PA-mm 写成这段精度。不把 MyEgo 或 EgoAVU 的数字写成本段成绩。
Not sold as VLA-ready. OpenXR / Vision hands are not robot actions. SGBM distance is not depth-camera truth. Paper PA-mm is not this take’s accuracy. MyEgo or EgoAVU figures are not this take’s score.
旁路:世界模型可读 E2 长双目 + IMU 做视频预测;短程 VLN 表征优先 E6(有头部位姿)。两者都不是本页主推,字段与 VLM 共用,任务标签分开写。
Side paths: world models may use long E2 stereo + IMU for video prediction; short VLN representation prefers E6 (head pose present). Neither is the primary pitch here — fields are shared, task labels stay separate.
文献对照 · 2026-09-21Literature check · 2026-09-21
下面是文献要求的输入形式,对照本页已经能写出的字段。不是这些模型在 E2 / E6 上的分数。
These are input forms from the papers, matched to fields this page can already emit. They are not model scores on E2 / E6.
Qwen3-VL 在每个视频时间块前写文本时间,例如 <3.0 seconds> 或 00:00:22,时间不只藏在位置编码里。本页用 time_sync 和关键帧写记录时间。播放器 0 秒仍是记录 22 秒,两套时钟都留。Qwen3-VL prefixes each video time patch with a text time such as <3.0 seconds> or 00:00:22. Time is not left only in positional encoding. This page writes record time from time_sync and keyframes. Player 0 s remains record 22 s; both clocks stay.
Qwen3-VL Technical Report, arXiv:2511.21631 (2025-11).Qwen3-VL Technical Report, arXiv:2511.21631 (2025-11).
Qwen3-VL 使用动态原生分辨率。MyEgo 把帧抽到提问时刻,默认大约 1 fps、32 帧;加帧并不稳定更好。训练读 E2 1600×1200 @ 30 Hz,E6 RGB 约 60 Hz。抽几帧是训练选择。4 Hz / 10 Hz 预览不进训练列。Qwen3-VL uses dynamic native resolution. MyEgo samples frames up to the question time, about 1 fps and 32 frames by default; more frames did not reliably help. Train on E2 1600×1200 @ 30 Hz and E6 RGB ~60 Hz. Frame count is a training choice. The 4 Hz / 10 Hz preview is not a training column.
MyEgo, CVPR 2026, pp. 40537–40547.MyEgo, CVPR 2026, pp. 40537–40547.
EgoAVU 与 EgoSound 写明:现有模型偏视觉,常对不上声音来自哪一块画面。E2 音频是 16 kHz 立体声,比左目曝光晚 97 ms。这个偏移写进样本。EgoAVU and EgoSound report that current models lean on vision and often miss which sound belongs to which visible source. E2 audio is 16 kHz stereo, 97 ms after the left-eye exposure. Write that offset on the sample.
EgoAVU, CVPR 2026, pp. 15805–15814. EgoSound, CVPR 2026.EgoAVU, CVPR 2026, pp. 15805–15814. EgoSound, CVPR 2026.
MyEgo 问的是佩戴者本人。该文开放问答里 GPT-5 约 46.1%,Qwen3-VL-8B-Instruct 约 36.4%,人工约 84.7%。模型会把佩戴者和其他人混在一起。本页 human 句子按佩戴者写,并且可核对。E6 为 32 秒,E2 为单段 682 秒,不是跨天记忆。这些分数不是本段成绩。MyEgo asks about the camera wearer. In that paper’s open QA, GPT-5 is about 46.1%, Qwen3-VL-8B-Instruct about 36.4%, and humans about 84.7%. Models mix the wearer with other people. Our human sentences stay wearer-scoped and checkable. E6 is 32 s; E2 is one 682 s episode, not multi-day memory. Those scores are not this take’s score.
同文还点名 InternVL3、InternVL3.5、Gemini-2.5 Pro。本页不跑这些模型。The same paper also names InternVL3, InternVL3.5, and Gemini-2.5 Pro. This page does not run those models.
分级数据Graded data
用 TOBI 已落地的命名:device → indexes → human → estimate。对应采购时可按深度选档,不空喊未交付能力。样例字段 = 交付字段。
Named with what TOBI already ships: device → indexes → human → estimate. Buyers pick depth; we do not claim undelivered layers. Sample fields = delivery fields.
视频与 IMU / 位姿。原生帧率原片,标定随序列号。
Video plus IMU / pose. Native-Hz source; calibration by serial.
标准化结构、多模态时间对齐、validity 质检。
Standard structure, multimodal time sync, validity QA.
让模型读佩戴者句子与关键帧:场景、物体、动作时间段。
Wearer sentences and keyframes: scene, objects, action spans.
手点、轨迹辅助、隐私打码。带 method,不当设备测量,不作机器人动作。
Hand points, track aids, privacy mosaic. Method-tagged; not device truth; not robot actions.
原片 + 叙述 + 文本时间戳即可开训;估计层可选。
Source + narration + text timestamps train; estimate optional.
交付选档Delivery pick文件、结构化结果、索引与 validity 报告一并给出。
Files, structured results, indexes and validity reports together.
STOCK按场景加购小时;默认未标注原片,可加叙述 / 估计档。
Buy hours by scene; default unlabeled source, optional narration / estimate.
技术方案Method
本机没有 CUDA,所以现在不跑 MANO。E2 用 Apple Vision 在原片上写 21 点,每个点带置信度。E6 的手仍是设备 OpenXR,不用这张图替换。
This machine has no CUDA, so MANO is not run. On E2, Apple Vision writes 21 points on the source frame, each with a confidence. E6 hands stay device OpenXR. This picture does not replace them.
原片上的 21 点进入 estimate,并写明 method。低于 0.5 的点标成无效,不删掉。
The 21 points go into estimate with a method field. Points below 0.5 are marked invalid, not deleted.
再加 WiLoR 的相机系 MANO,仍放在 estimate。要世界轨迹时才用 HaWoR,尺度用本段双目。
Add WiLoR camera-frame MANO, still under estimate. Use HaWoR only for a world trajectory, and take scale from this episode’s stereo.
不把两目视差写成米。不把 E6 的设备关节换成视觉网格。不把对齐后的论文毫米数写成这段的精度。
Do not turn two-eye disparity into meters. Do not replace E6 device joints with a vision mesh. Do not treat a paper’s aligned millimeter score as this take’s accuracy.
这一帧来自已公开的 E2 样例左目,文件名序号 0120,曝光时刻没有重新读取。另一段原片(uid 532f0cf9a78d,不是 20260917)已按同样方法在左右目各跑 613 秒:左目 593 秒有手,右目 588 秒有手。那份 1 Hz 坐标不替代 30 Hz 原片。
This frame is the published E2 left eye, filename index 0120. The exposure time was not re-read. A different take (uid 532f0cf9a78d, not 20260917) was run the same way for 613 seconds on each eye: a hand is present on 593 left seconds and 588 right seconds. That 1 Hz file does not replace the 30 Hz source.
信息架构Architecture
同一段采集拆成三层。缺的列留空。训练读原片原生帧率,不读 4 Hz / 10 Hz 预览。
One capture splits into three namespaces. Missing columns stay empty. Train at native source Hz, not the 4 Hz / 10 Hz preview.
device
只收采集设备或出厂标定写进文件的流。
Only streams written by the capture device or factory calibration.
estimate
必须带 method,并写 not_device_output。
Must carry method and not_device_output.
human
句子不改写成检测器输出。
Sentences are never rewritten as detector output.
embodied_layers.v1.json E2 validity E6 validity vision_pipeline_v1.json
层说明Layer notes
设备陀螺和加速度计给出角速度和线加速度。E2 陀螺和加速度约 799 Hz,磁力计 200 Hz,加速度模长抽样均值 9.83 m/s²。E6 陀螺和加速度约 1015 Hz。E2 加速度比左目曝光起点早 124 毫秒。距离片底部曲线来自 gyro.csv。
E2 这 682 秒按 5 秒窗口角速度均方根切开。中位数 0.352 rad/s。高于 1.8 倍且持续至少 10 秒的画成橙色。这是头戴角速度,不是手部动作,也不是位姿。最长日常段 150–505 秒;本段最高 595–605 秒,均方根 2.188 rad/s。积分航向会漂移,不是 SLAM,E2 也没有头部位姿文件。
Device gyro and accelerometer record angular velocity and linear acceleration. E2 gyro/accel ≈ 799 Hz, magnetometer 200 Hz, accel magnitude subsample averages 9.83 m/s². E6 ≈ 1015 Hz. E2 accel starts 124 ms before left-eye exposure. The curve under the distance clip is gyro.csv.
The 682 s are split by 5-second gyro RMS (median 0.352 rad/s). Spans above 1.8× that median and ≥10 s are orange. That is head-worn rate, not a hand verb or pose. Longest usual stretch 150–505 s; peak 595–605 s at RMS 2.188. Integrated heading drifts — not SLAM; E2 has no head-pose file.
橙:角速度升高。灰:其余。40–48 秒样例落在第二段橙色(30–100 秒,RMS 1.818)。 motion_spans.jsonl
Orange: elevated rate. Gray: the rest. The 40–48 s sample sits in the second orange span (30–100 s, RMS 1.818). motion_spans.jsonl
两段都没有深度相机文件。这 8 秒取自 E2 40.0–48.0 秒,4 Hz,只供核对。左为 Camera0 原片 1600×1200(未校正),屏幕与远处的人打了马赛克,佩戴者手不打。右为校正后 Z = f·B / d(四元数 xyzw):焦距约 253 px,基线约 66 mm。SGBM 后左右一致 + 置信门控:有效像素约 40%,中位距离约 1.34 m,有效区置信中位约 0.90,裁在 0.25–6 m。黑块不填洞。本机无 ximgproc,未跑 WLS。
estimate 层输出。径向只用前三个系数,未经棋盘格复核。音频 16 kHz 立体声,比左目曝光晚 97 ms。训练读 30 Hz 原片。FoundationStereo(CVPR 2025,第三方)槽位已留,无 CUDA 未跑。E6 RGB 基线约 60 mm。标称 SLAM 5 mm 需按任务验证,本段无地面真值。
Neither episode has a depth-camera file. These 8 s are E2 40.0–48.0 at 4 Hz for checking. Left is original Camera0 1600×1200 (not rectified); screen and distant people mosaicked; wearer hands kept. Right is rectified Z = f·B / d (xyzw): ~253 px focal, ~66 mm baseline. After SGBM: LR check + confidence gate → ~40% valid, median ~1.34 m, confidence median ~0.90, clipped 0.25–6 m. Black not filled. No ximgproc/WLS on this host.
Estimate-layer only. Three radial coeffs, no checkerboard check. Audio 16 kHz stereo, +97 ms. Train at 30 Hz source. FoundationStereo reserved, not run (no CUDA). E6 RGB baseline ~60 mm. Nominal 5 mm SLAM still needs per-task verification; no ground truth here.
回到样例Back to sample depth_sample.json stereo_distance_sample.json E2 全长样例Full E2 sample
E2 标注片已画:Vision 21 点手、人脸、远处人体、亮屏幕、蓝色墙标、置信度 ≥ 0.40 的文字。707 秒里:右手 525、左手 440、双手 410、人体 221、人脸 77、蓝标 32、亮屏 642。无类别名。OCR 常误读,弱结果不当招牌。手部包装 hand_2d_vision21.json(estimate,非 OpenXR,关节置信空,HaMeR 未跑)。
无实例掩码。E6 物体只有人工时间段;E6 手为设备 OpenXR 26 关节 hand_openxr_device.json(文件在 estimate 目录,命名空间仍是 device)。类别框/掩码另开文件,不改写现有框。多目三角化、4 目以上离群剔除、关节 Butterworth 不用于这两段:E2 只有两目,且 707 行只有计数、没有关节点坐标。E6 手部索引按 RGB 帧对齐,最大间隔约 19.6 秒,不能平滑补洞。
E2 annotated preview draws Vision 21-point hands, faces, distant people, bright screens, blue wall marks, and text at confidence ≥0.40. Across 707 s: right 525, left 440, both 410, person 221, face 77, blue mark 32, screen 642. No class names. OCR often misreads; weak hits are not treated as signage. Hand package hand_2d_vision21.json (estimate, not OpenXR; joint confidence empty; HaMeR not run).
No instance masks. E6 objects are human time spans only. E6 hands are device OpenXR 26 joints: hand_openxr_device.json (folder is historical; namespace stays device). Class boxes or masks go in a separate file; existing boxes are not rewritten. Multi-camera triangulation, 4-plus-view outlier rejection, and joint Butterworth are not used on these takes: E2 has two eyes, and the 707 rows are counts without joint coordinates. The E6 hand index follows the RGB clock and has a gap of about 19.6 s, so the hole is not filled by smoothing.
一段采集先对齐时钟,再按算法各写一份。已落盘:E6 设备关节与头部位姿;E2 手部 21 点、打码、陀螺窗口,以及页首 8 秒距离样例。场景句子仍是人工核对。预览片只用来看,训练读原始视频。
Align the clock first, then each algorithm writes its own file. On disk: E6 device joints and head pose; E2 21-point hands, mosaic, gyro windows, and the 8 s distance sample above. Scene sentences stay human. Preview is for looking; training reads source video.
交互样例Interactive sample
点步骤或用左右键:两路同步跳转并播放,只响原始片。播放器 0 秒 = 记录 22 秒。框、掩码、6DoF、接触力、关节置信度为空。
Click a step or use the arrows: both players jump and play together; only raw stays audible. Player 0 s = record 22 s. Boxes, masks, 6DoF, contact force and joint confidence stay empty.
E6 · 22.00–23.17 · RGB 1320–1390
双手在桌面操作区,握持深色手柄。左手追踪还没稳定。
Both hands at the desk, gripping dark controllers. Left tracking is not stable yet.
拇指到食指是关节几何距离,不是接触力。框 / 掩码 / 6DoF / 关节置信度为空。
Thumb–index is joint geometry, not contact force. Boxes / masks / 6DoF / joint confidence stay empty.
E6 · 23.17–25.50
左右手同时有效,仍握持深色手柄,看向桌面终端。
Both hands stay active, still gripping the controllers and facing the terminal.
拇指–食指:左 73 mm,右 56 mm。不是接触力。
Thumb–index: left 73 mm, right 56 mm. Not contact force.
E6 · 25.50–27.52
转到纸卷和布。右手更稳,左手时有时无。物体没有框。
Hands move to a paper roll and cloth. Right is steadier; left comes and goes. No object boxes.
E6 · 27.52–29.50
左右手同时有效,继续操作纸卷和布。头几乎不位移。
Both hands stay active on the paper and cloth. Head barely nets displacement.
E6 · 29.50–31.20
双手捏持小零件。拇指到食指更近。没有物体框。
Both hands hold a small part. Thumb–index is closer. No object box.
24 mm 是关节几何,不是夹持力。
24 mm is joint geometry, not grip force.
E6 · 31.20–结束E6 · 31.20–end
左手持手机看屏幕。右手有效 0%。手柄位姿文件仍是一行零。
Left holds a phone and looks at the screen. Right active 0%. Controller-pose file is one zero row.
屏幕文字未做检测级打码。
Screen text is not detector-masked.
E2 · 20260917_133640
11 分 22 秒。设备只有左右目、IMU、磁力计和 16 kHz 立体声。骨架、马赛克和右上视差不是设备输出。一起播放时只响你按下的那一路。
11 min 22 s. Device files are the two eyes, IMU, mag and 16 kHz stereo. Skeleton, mosaic and corner disparity are estimates. Only the player you press stays audible.
训练门禁Training gate
源文件不删。下面这些不写入训练列,也不用别的数填上。
Source files stay. These do not enter a training column, and nothing is invented to fill them.
对照Compare
预览 mp4 用来核对。训练读原始视频和索引。
Preview mp4 is for checking. Training reads source video and indexes.
| 格式 | E6 | E2 |
|---|---|---|
| 原始视频 | rgb / tracking / ctrl,60 Hz 量级 | Camera0、Camera1,1600×1200,30 Hz |
| 预览 | 10 Hz,从 22 秒起,带关节投影和声音 | 4 Hz 标注片,带声音;另有无标注并排片 |
| 时间索引 | time_sync.json · rgb_index.jsonl | time_sync.json · stereo_index.jsonl |
| 手 | hands_flags.jsonl + hand_openxr_device.json | hand_2d_vision21.json + e2-findings.jsonl |
| 句子 | narration.jsonl · e6_keyframe_step.v1.json | narration.jsonl |
| 物体 | object_spans.jsonl | 没有类别框 |
| 有效性 | e6 validity | e2 validity |
| Format | E6 | E2 |
|---|---|---|
| Source video | rgb / tracking / ctrl, about 60 Hz | Camera0 and Camera1, 1600×1200, 30 Hz |
| Preview | 10 Hz from record 22 s, with joint projection and audio | 4 Hz annotated preview with audio, plus an unannotated side-by-side |
| Time index | time_sync.json · rgb_index.jsonl | time_sync.json · stereo_index.jsonl |
| Hands | hands_flags.jsonl + hand_openxr_device.json | hand_2d_vision21.json + e2-findings.jsonl |
| Sentences | narration.jsonl · e6_keyframe_step.v1.json | narration.jsonl |
| Objects | object_spans.jsonl | No class boxes |
| Validity | e6 validity | e2 validity |
客户要哪一种,就下哪一种文件,不要把估计写成设备测量。
Take the file that matches the need. Do not write an estimate as a device measurement.
| 要训练的输入 | 用 E6 | 用 E2 |
|---|---|---|
| 手怎么动 | 设备 OpenXR 26 关节,约 19 Hz | Vision 21 点,估计 |
| 头在哪里 | head_pose.csv,60 Hz | 无头部位姿;陀螺积分会漂移 |
| 场景一句 | 四句人工叙述 | 一句人工叙述 |
| 文字和隐私 | 本页预览无检测级打码 | 人脸、远处人体、文字 ≥0.40、亮屏、蓝标 |
| 远近 | 预览无视差;SLAM 5 mm 标称 | SGBM + LR + 置信;FoundationStereo 未跑 |
| 声音 | 44.1 kHz 单声道,+288 ms | 16 kHz 立体声,+97 ms |
| Input to train | Use E6 | Use E2 |
|---|---|---|
| How the hands move | Device OpenXR 26 joints, about 19 Hz | Vision 21-point estimate |
| Where the head is | head_pose.csv, 60 Hz | No head pose; integrated gyro drifts |
| One scene sentence | Four human sentences | One human sentence |
| Text and privacy | This preview is not detector-masked | Faces, distant people, text at 0.40+, bright screens, blue wall marks |
| Near and far | No disparity on the preview; the 5 mm SLAM figure is nominal | SGBM + left-right check + confidence gate; FoundationStereo not run |
| Audio | 44.1 kHz mono, +288 ms | 16 kHz stereo, +97 ms |
两段都没有接触力、物体 6DoF、机器人末端动作。E6 的 grayscale tracking 与朝下 ctrl 在观测里;ctrl 不是手柄位姿(手柄文件一行零)。E2 没有这些灰度相机。
Neither episode has contact force, object 6DoF, or robot eef action. E6 grayscale tracking and downward ctrl are observations; ctrl is not controller pose. E2 has no grayscale cameras.
07 · 交付样例07 · Delivery samples
主推 VLM 包落在 device / human / indexes;estimate 可选。MCAP 与 LeRobot 目录已预留,当前未导出。下面四类是真实交付清单,不是口号。
Primary VLM packs land as device / human / indexes; estimate optional. MCAP and LeRobot folders are reserved, not exported yet. The four cards below are the real delivery checklist.
根据项目选深度。默认 D2;VLM 主推加 D3;手部辅助加 D4。
Pick depth per project. Default D2; add D3 for VLM; add D4 for hand aids.
01
02
03
04
VLM 最小集:device/ 原片、标定与音频偏移 · human/narration.jsonl(及 E6 关键帧)· indexes/time_sync.json · validity。预览 mp4 与未门控的估计距离不进训练列,公开模型分数也不写入。
Minimum VLM set: device/ source, calibration, and the audio offset · human/narration.jsonl (plus E6 keyframes) · indexes/time_sync.json · validity. Do not put preview mp4, ungated estimated distance, or public model scores into a training column.
episode_bundle.v1.json embodied_layers.v1.json 交付分级说明Delivery tier note
开始构建Get started
查看 E2 / E6 / E8 规格与量产状态。Review E2 / E6 / E8 specs and volume status.
02 · 现货Stock按场景加购小时,提交报价。Add hours by scene and request a quote.
03 · 分级GradesD1→D4:原片 / 对齐 / 语义 / 估计。D1→D4: source / align / semantic / estimate.