具身数据Embodied Data English

TOBI

主推 VLM:一次采集,拆成可训的视觉–语言层。Built for VLM: one capture, split into trainable vision–language layers.

原生帧率原片、文本时间戳、带偏移的音频、佩戴者视角的人工句子。设备测量与后处理估计分开。这是人佩戴观测,不是机器人动作包。

Native-Hz source, text timestamps, audio with its offset, and wearer-scoped human sentences. Device measurements and later estimates stay apart. Human-worn observation — not a robot action pack.

VLM · 主推primary E6 · 32.0 秒 · 六路32.0 s · six cameras E2 · 682.17 秒 · 双目682.17 s · stereo
E2 left eye with Vision 21 hand points
E2 · estimate 左目样例。青 = 左手,橙 = 右手,灰 = 置信度低于 0.5。屏幕已打码,佩戴者的手保留。 Left-eye sample. Cyan = left hand, orange = right hand, gray = confidence below 0.5. The screen is mosaicked; the wearer’s hands stay.
E2 · estimate · 40.0–48.0 s · 4 Hz 左:原片左目。右:SGBM + 左右一致 + 置信门控。黑块 = 低置信 / 匹配失败。不是深度相机。 Left: original Camera0. Right: SGBM + LR check + confidence gate. Black = low confidence. Not a depth camera.
预览只供核对。训练读 Camera0 / Camera1 原片 1600×1200 @ 30 Hz。有效像素约 40%,距离中位约 1.34 m。 Preview for checking only. Train on Camera0 / Camera1 at 1600×1200 @ 30 Hz. ~40% valid pixels, median distance ~1.34 m.

主推方案Primary solution

VLM:第一人称观测 → 可训的视觉–语言层

VLM: egocentric observation → trainable vision–language layers

第一人称 VLM 主要读四样:原生帧率画面、文本时间戳、带偏移的音频、佩戴者视角的人工句子。E6 / E2 按 device / human / indexes 交付这四样。手轨迹与 IMU 可作条件,不写成机器人动作。世界模型与短程 VLN 仍是旁路;标准 VLA 仍缺末端与接触力。

Egocentric VLMs mainly read four things: native-Hz frames, text timestamps, audio with its offset, and wearer-scoped human sentences. E6 / E2 deliver those four under device / human / indexes. Hand tracks and IMU may condition a model. They are not robot actions. World models and short VLN stay side paths; standard VLA still needs an end-effector and contact force.

01 · 主推01 · PRIMARY

VLM

适配强。原片 + 叙述即可开训 / 评测;手与检测为可选辅助。

Strong fit. Source + narration is enough to train or eval; hands and detectors stay optional helpers.

02

世界模型

World model

可部分用。长双目与多相机适合视频预测;缺可交互环境状态。

Partial. Long stereo and multi-cam suit video prediction; no interactive world state.

03

VLN

可部分用(表征)。E6 有头部位姿;无地图与标准 path–指令对。

Partial (representation). E6 has head pose; no map or standard path–instruction pairs.

04

VLA

当前不适配。无末端、夹爪、接触力;手柄位姿为零行。

Not a fit today. No eef, gripper, or contact force; controller pose is a zero row.

VLM 训练包读什么

What a VLM pack reads

  1. device
    原片视频(原生帧率) Source video (native Hz) E2 Camera0/1 · 1600×1200 @ 30 Hz。E6 RGB 约 60 Hz。训练不读 4 Hz / 10 Hz 预览。 E2 Camera0/1 · 1600×1200 @ 30 Hz. E6 RGB ~60 Hz. Do not train on 4 Hz / 10 Hz preview.
  2. audio
    音频(写明偏移) Audio (offset written) E2 为 16 kHz 立体声,比左目曝光晚 97 ms。偏移写进样本。不写成已经完成的音画对应。 E2 is 16 kHz stereo, 97 ms after the left-eye exposure. Write that offset on the sample. Do not call it solved audio-visual grounding.
  3. human
    叙述与关键帧句 Narration and keyframe sentences E6 四句 + 六步关键帧;E2 一句。句子不改写成检测器输出。 E6: four sentences + six keyframes; E2: one sentence. Never rewrite a sentence as detector output.
  4. estimate · 可选 estimate · optional
    手点 / 马赛克 / 检测计数 Hand points / mosaic / detection counts Vision 21 点、打码、OCR 弱结果:带 method,不当设备测量,可作 grounding 辅助。 Vision 21, mosaic, weak OCR: carry method, not device measurement; optional grounding aids.
  5. indexes
    时间同步与 validity Time sync and validity time_sync、rgb/stereo 索引、validity 报告。缺列留空,不填假值。 time_sync, rgb/stereo indexes, validity reports. Missing columns stay empty — no invented fills.
  1. 01

    现在就能交付

    Ship now

    第一人称 VQA、场景描述、手–物理解:视觉读 device 原片,语言读 human,时间读 indexes。工作台六步句可直接核对。

    Egocentric VQA, scene captioning, hand–object understanding: vision from device source, language from human, time from indexes. The six workbench sentences are checkable as-is.

  2. 02

    怎么组一条样本

    How one sample is built

    对齐时钟 → 按索引取原生帧 → 挂叙述时间段 → 写成文本时间戳并写上音频偏移 → 可选挂手点或隐私掩码。预览 mp4 只给人看。

    Align clocks → seek native frames by index → attach narration spans → write a text timestamp and the audio offset → optionally attach hand points or privacy masks. Preview mp4 is for humans only.

  3. 03

    明确不承诺

    Not claimed

    不写成 VLA-ready。不把 OpenXR / Vision 手当成 robot action。不把 SGBM 距离当成深度相机真值。不把论文 PA-mm 写成这段精度。不把 MyEgo 或 EgoAVU 的数字写成本段成绩。

    Not sold as VLA-ready. OpenXR / Vision hands are not robot actions. SGBM distance is not depth-camera truth. Paper PA-mm is not this take’s accuracy. MyEgo or EgoAVU figures are not this take’s score.

旁路:世界模型可读 E2 长双目 + IMU 做视频预测;短程 VLN 表征优先 E6(有头部位姿)。两者都不是本页主推,字段与 VLM 共用,任务标签分开写。

Side paths: world models may use long E2 stereo + IMU for video prediction; short VLN representation prefers E6 (head pose present). Neither is the primary pitch here — fields are shared, task labels stay separate.

文献对照 · 2026-09-21Literature check · 2026-09-21

公开文献里,第一人称 VLM 在读的四件事

Four inputs public egocentric VLM papers actually read

下面是文献要求的输入形式,对照本页已经能写出的字段。不是这些模型在 E2 / E6 上的分数。

These are input forms from the papers, matched to fields this page can already emit. They are not model scores on E2 / E6.

Qwen3-VL

文本时间戳Text timestamps

Qwen3-VL 在每个视频时间块前写文本时间,例如 <3.0 seconds> 或 00:00:22,时间不只藏在位置编码里。本页用 time_sync 和关键帧写记录时间。播放器 0 秒仍是记录 22 秒,两套时钟都留。Qwen3-VL prefixes each video time patch with a text time such as <3.0 seconds> or 00:00:22. Time is not left only in positional encoding. This page writes record time from time_sync and keyframes. Player 0 s remains record 22 s; both clocks stay.

Qwen3-VL Technical Report, arXiv:2511.21631 (2025-11).Qwen3-VL Technical Report, arXiv:2511.21631 (2025-11).

MyEgo

原生帧,抽帧留到训练时Native frames; subsample at train time

Qwen3-VL 使用动态原生分辨率。MyEgo 把帧抽到提问时刻,默认大约 1 fps、32 帧;加帧并不稳定更好。训练读 E2 1600×1200 @ 30 Hz,E6 RGB 约 60 Hz。抽几帧是训练选择。4 Hz / 10 Hz 预览不进训练列。Qwen3-VL uses dynamic native resolution. MyEgo samples frames up to the question time, about 1 fps and 32 frames by default; more frames did not reliably help. Train on E2 1600×1200 @ 30 Hz and E6 RGB ~60 Hz. Frame count is a training choice. The 4 Hz / 10 Hz preview is not a training column.

MyEgo, CVPR 2026, pp. 40537–40547.MyEgo, CVPR 2026, pp. 40537–40547.

EgoAVU

声音和画面一起留,并写偏移Keep audio with the picture, and write the offset

EgoAVU 与 EgoSound 写明:现有模型偏视觉,常对不上声音来自哪一块画面。E2 音频是 16 kHz 立体声,比左目曝光晚 97 ms。这个偏移写进样本。EgoAVU and EgoSound report that current models lean on vision and often miss which sound belongs to which visible source. E2 audio is 16 kHz stereo, 97 ms after the left-eye exposure. Write that offset on the sample.

EgoAVU, CVPR 2026, pp. 15805–15814. EgoSound, CVPR 2026.EgoAVU, CVPR 2026, pp. 15805–15814. EgoSound, CVPR 2026.

MyEgo

佩戴者的句子,不是场景摘要Wearer sentences, not a scene summary

MyEgo 问的是佩戴者本人。该文开放问答里 GPT-5 约 46.1%,Qwen3-VL-8B-Instruct 约 36.4%,人工约 84.7%。模型会把佩戴者和其他人混在一起。本页 human 句子按佩戴者写,并且可核对。E6 为 32 秒,E2 为单段 682 秒,不是跨天记忆。这些分数不是本段成绩。MyEgo asks about the camera wearer. In that paper’s open QA, GPT-5 is about 46.1%, Qwen3-VL-8B-Instruct about 36.4%, and humans about 84.7%. Models mix the wearer with other people. Our human sentences stay wearer-scoped and checkable. E6 is 32 s; E2 is one 682 s episode, not multi-day memory. Those scores are not this take’s score.

同文还点名 InternVL3、InternVL3.5、Gemini-2.5 Pro。本页不跑这些模型。The same paper also names InternVL3, InternVL3.5, and Gemini-2.5 Pro. This page does not run those models.

分级数据Graded data

完整管线:真实佩戴记录 → 按深度加工成可训资产

One pipeline: worn capture → graded assets for training

用 TOBI 已落地的命名:device → indexes → human → estimate。对应采购时可按深度选档,不空喊未交付能力。样例字段 = 交付字段。

Named with what TOBI already ships: device → indexes → human → estimate. Buyers pick depth; we do not claim undelivered layers. Sample fields = delivery fields.

E6 worn capture
D1 · DEVICE

原片数据

Source data

视频与 IMU / 位姿。原生帧率原片,标定随序列号。

Video plus IMU / pose. Native-Hz source; calibration by serial.

  • E2 双目 / E6 多路 RGB
  • IMU · 音频偏移
  • 相机内外参
  • E2 stereo / E6 multi RGB
  • IMU · audio offset
  • Intrinsics / extrinsics
E6 synced frames
D2 · INDEXES

基础对齐

Aligned base

标准化结构、多模态时间对齐、validity 质检。

Standard structure, multimodal time sync, validity QA.

  • time_sync.json
  • rgb / stereo 索引
  • 缺列留空
  • time_sync.json
  • rgb / stereo indexes
  • Empty columns stay empty
Workbench narration
D3 · HUMAN

语义数据

Semantic data

让模型读佩戴者句子与关键帧:场景、物体、动作时间段。

Wearer sentences and keyframes: scene, objects, action spans.

  • narration.jsonl
  • 关键帧六步 / 物体时间段
  • 人工核对,非检测器改写
  • narration.jsonl
  • Six keyframes / object spans
  • Human check, not detector rewrite
Hand estimate points
D4 · ESTIMATE

行为估计(可选)

Behavior estimate (optional)

手点、轨迹辅助、隐私打码。带 method,不当设备测量,不作机器人动作。

Hand points, track aids, privacy mosaic. Method-tagged; not device truth; not robot actions.

  • Vision 21 / OpenXR 26
  • 马赛克 · 弱 OCR
  • SGBM 距离样例(非深度真值)
  • Vision 21 / OpenXR 26
  • Mosaic · weak OCR
  • SGBM distance sample (not depth truth)
D1
原片Source 双目 / 多路视频 · IMU · 标定 Stereo / multi video · IMU · calibration
清洗对齐
质检
clean · sync
QA
D2
对齐Aligned 标准化结构 · 多模态同步 · 场景索引 Standard structure · multimodal sync · scene index
叙述挂接
人工复核
attach narration
human review
D3
语义Semantic 任务语义 · 物体 / 动作 · 结构化句子 Task semantics · object / action · structured sentences
估计提取
写 method
estimate extract
write method
D4
估计Estimate 手部关键点 · 轨迹辅助 · 行为理解辅助 Hand keypoints · track aids · behavior helpers

技术方案Method

手部识别:先写能核对的 21 点

Hands: write 21 points you can check

本机没有 CUDA,所以现在不跑 MANO。E2 用 Apple Vision 在原片上写 21 点,每个点带置信度。E6 的手仍是设备 OpenXR,不用这张图替换。

This machine has no CUDA, so MANO is not run. On E2, Apple Vision writes 21 points on the source frame, each with a confidence. E6 hands stay device OpenXR. This picture does not replace them.

E2 left-eye Vision 21 points
E2 左目。青 = 左手,橙 = 右手,灰 = 置信度低于 0.5。屏幕和远处的人已打码,佩戴者的手保留。 E2 left eye. Cyan = left hand, orange = right hand, gray = confidence below 0.5. Screens and distant people are mosaicked. The wearer’s hands stay visible.
  1. 01

    现在

    Now

    原片上的 21 点进入 estimate,并写明 method。低于 0.5 的点标成无效,不删掉。

    The 21 points go into estimate with a method field. Points below 0.5 are marked invalid, not deleted.

  2. 02

    有 NVIDIA 之后

    When NVIDIA is available

    再加 WiLoR 的相机系 MANO,仍放在 estimate。要世界轨迹时才用 HaWoR,尺度用本段双目。

    Add WiLoR camera-frame MANO, still under estimate. Use HaWoR only for a world trajectory, and take scale from this episode’s stereo.

  3. 03

    不做

    Not done

    不把两目视差写成米。不把 E6 的设备关节换成视觉网格。不把对齐后的论文毫米数写成这段的精度。

    Do not turn two-eye disparity into meters. Do not replace E6 device joints with a vision mesh. Do not treat a paper’s aligned millimeter score as this take’s accuracy.

18/21左手有效点Left joints kept
15/21右手有效点Right joints kept
0.5置信度门限Confidence gate
estimate不是设备输出Not a device output

这一帧来自已公开的 E2 样例左目,文件名序号 0120,曝光时刻没有重新读取。另一段原片(uid 532f0cf9a78d,不是 20260917)已按同样方法在左右目各跑 613 秒:左目 593 秒有手,右目 588 秒有手。那份 1 Hz 坐标不替代 30 Hz 原片。

This frame is the published E2 left eye, filename index 0120. The exposure time was not re-read. A different take (uid 532f0cf9a78d, not 20260917) was run the same way for 613 seconds on each eye: a hand is present on 593 left seconds and 588 right seconds. That 1 Hz file does not replace the 30 Hz source.

信息架构Architecture

device · estimate · human

同一段采集拆成三层。缺的列留空。训练读原片原生帧率,不读 4 Hz / 10 Hz 预览。

One capture splits into three namespaces. Missing columns stay empty. Train at native source Hz, not the 4 Hz / 10 Hz preview.

device

设备写过的

From the device

只收采集设备或出厂标定写进文件的流。

Only streams written by the capture device or factory calibration.

  • RGB / 灰度 / IMU / 磁力计 / 音频
  • 头部位姿 · OpenXR 手
  • 相机内外参 · 时间同步
  • RGB / grayscale / IMU / magnetometer / audio
  • Head pose · OpenXR hands
  • Camera intrinsics & extrinsics · time sync

estimate

离线算法

Offline algorithms

必须带 method,并写 not_device_output。

Must carry method and not_device_output.

  • Vision 21 点手 · SGBM 距离
  • 马赛克 · OCR
  • MediaPipe / RTMPose / HaMeR / FoundationStereo 槽位保留,未在原片重跑
  • Vision 21-point hands · SGBM distance
  • Mosaic · OCR
  • MediaPipe / RTMPose / HaMeR / FoundationStereo slots reserved, not re-run on source

human

人工核对

Human review

句子不改写成检测器输出。

Sentences are never rewritten as detector output.

  • 叙述 · 六步关键帧句
  • 物体时间段
  • 接受标记
  • Narration · six keyframe sentences
  • Object time spans
  • Accept flags

embodied_layers.v1.json E2 validity E6 validity vision_pipeline_v1.json

层说明Layer notes

IMU · 距离 · 检测

IMU · Distance · Detection

设备测量Device不是从画面反推Not from the picture

IMU 数据

IMU data

设备陀螺和加速度计给出角速度和线加速度。E2 陀螺和加速度约 799 Hz,磁力计 200 Hz,加速度模长抽样均值 9.83 m/s²。E6 陀螺和加速度约 1015 Hz。E2 加速度比左目曝光起点早 124 毫秒。距离片底部曲线来自 gyro.csv。

E2 这 682 秒按 5 秒窗口角速度均方根切开。中位数 0.352 rad/s。高于 1.8 倍且持续至少 10 秒的画成橙色。这是头戴角速度,不是手部动作,也不是位姿。最长日常段 150–505 秒;本段最高 595–605 秒,均方根 2.188 rad/s。积分航向会漂移,不是 SLAM,E2 也没有头部位姿文件。

Device gyro and accelerometer record angular velocity and linear acceleration. E2 gyro/accel ≈ 799 Hz, magnetometer 200 Hz, accel magnitude subsample averages 9.83 m/s². E6 ≈ 1015 Hz. E2 accel starts 124 ms before left-eye exposure. The curve under the distance clip is gyro.csv.

The 682 s are split by 5-second gyro RMS (median 0.352 rad/s). Spans above 1.8× that median and ≥10 s are orange. That is head-worn rate, not a hand verb or pose. Longest usual stretch 150–505 s; peak 595–605 s at RMS 2.188. Integrated heading drifts — not SLAM; E2 has no head-pose file.

橙:角速度升高。灰:其余。40–48 秒样例落在第二段橙色(30–100 秒,RMS 1.818)。 motion_spans.jsonl

Orange: elevated rate. Gray: the rest. The 40–48 s sample sits in the second orange span (30–100 s, RMS 1.818). motion_spans.jsonl

交互样例Interactive sample

标注工作台

Annotation workbench

点步骤或用左右键:两路同步跳转并播放,只响原始片。播放器 0 秒 = 记录 22 秒。框、掩码、6DoF、接触力、关节置信度为空。

Click a step or use the arrows: both players jump and play together; only raw stays audible. Player 0 s = record 22 s. Boxes, masks, 6DoF, contact force and joint confidence stay empty.

1 / 6 步Step 1 / 6 Source 22.00 s 预览Preview 0.00 s
E6 RAW
E6 ANNOTATED

E2 · 20260917_133640

双目上的后处理

Estimates on stereo

11 分 22 秒。设备只有左右目、IMU、磁力计和 16 kHz 立体声。骨架、马赛克和右上视差不是设备输出。一起播放时只响你按下的那一路。

11 min 22 s. Device files are the two eyes, IMU, mag and 16 kHz stereo. Skeleton, mosaic and corner disparity are estimates. Only the player you press stays audible.

E2 RAW
E2 ANNOTATED

打开 E2 样例说明E2 sample notes

训练门禁Training gate

清理:这些不进训练列

Drop what should not train

源文件不删。下面这些不写入训练列,也不用别的数填上。

Source files stay. These do not enter a training column, and nothing is invented to fill them.

  1. 用 4 Hz / 10 Hz 预览代替 E2 30 Hz、E6 ~60 Hz 原片。
  2. 把后处理估计写成设备测量。
  3. 文字置信度 < 0.40,或不足两个字母/数字。
  4. 盖住佩戴者双手的人体框。
  5. 深度、物体 6DoF、关节置信度缺失时补假值。
  6. E6 里 active ≠ 1 的关节当作可见。
  7. 没有头部位姿却套用工装时间对齐(这两段都没有套用)。
  8. 把人工句子写成检测器输出。
  9. 双目匹配失败像素插值补洞。
  10. 用两目去跑至少三目的三角化,或在没有左右对应关节点时写出 3D 手。
  11. 把手部索引里的时间空洞用平滑填上,再当成识别轨迹。
  1. Train on 4 Hz / 10 Hz preview instead of native source Hz.
  2. Write a later estimate as a device measurement.
  3. Drop text below confidence 0.40, or with fewer than two letters or digits.
  4. Person boxes that cover the wearer’s hands.
  5. Invent depth, object 6DoF, or joint confidence when absent.
  6. Treat E6 joints with active ≠ 1 as visible.
  7. Apply factory time alignment without head pose to check.
  8. Treat a human sentence as detector output.
  9. Fill stereo match failures with interpolation.
  10. Run a three-or-more-camera triangulation on two eyes, or write a 3D hand without left-right joint correspondence.
  11. Smooth across a gap in the hand index and call the result a recognition track.

对照Compare

按输出格式拆

By output format

预览 mp4 用来核对。训练读原始视频和索引。

Preview mp4 is for checking. Training reads source video and indexes.

格式E6E2
原始视频rgb / tracking / ctrl,60 Hz 量级Camera0、Camera1,1600×1200,30 Hz
预览10 Hz,从 22 秒起,带关节投影和声音4 Hz 标注片,带声音;另有无标注并排片
时间索引time_sync.json · rgb_index.jsonltime_sync.json · stereo_index.jsonl
hands_flags.jsonl + hand_openxr_device.jsonhand_2d_vision21.json + e2-findings.jsonl
句子narration.jsonl · e6_keyframe_step.v1.jsonnarration.jsonl
物体object_spans.jsonl没有类别框
有效性e6 validitye2 validity
FormatE6E2
Source videorgb / tracking / ctrl, about 60 HzCamera0 and Camera1, 1600×1200, 30 Hz
Preview10 Hz from record 22 s, with joint projection and audio4 Hz annotated preview with audio, plus an unannotated side-by-side
Time indextime_sync.json · rgb_index.jsonltime_sync.json · stereo_index.jsonl
Handshands_flags.jsonl + hand_openxr_device.jsonhand_2d_vision21.json + e2-findings.jsonl
Sentencesnarration.jsonl · e6_keyframe_step.v1.jsonnarration.jsonl
Objectsobject_spans.jsonlNo class boxes
Validitye6 validitye2 validity

按算法需求拆

By algorithm need

客户要哪一种,就下哪一种文件,不要把估计写成设备测量。

Take the file that matches the need. Do not write an estimate as a device measurement.

要训练的输入用 E6用 E2
手怎么动设备 OpenXR 26 关节,约 19 HzVision 21 点,估计
头在哪里head_pose.csv,60 Hz无头部位姿;陀螺积分会漂移
场景一句四句人工叙述一句人工叙述
文字和隐私本页预览无检测级打码人脸、远处人体、文字 ≥0.40、亮屏、蓝标
远近预览无视差;SLAM 5 mm 标称SGBM + LR + 置信;FoundationStereo 未跑
声音44.1 kHz 单声道,+288 ms16 kHz 立体声,+97 ms
Input to trainUse E6Use E2
How the hands moveDevice OpenXR 26 joints, about 19 HzVision 21-point estimate
Where the head ishead_pose.csv, 60 HzNo head pose; integrated gyro drifts
One scene sentenceFour human sentencesOne human sentence
Text and privacyThis preview is not detector-maskedFaces, distant people, text at 0.40+, bright screens, blue wall marks
Near and farNo disparity on the preview; the 5 mm SLAM figure is nominalSGBM + left-right check + confidence gate; FoundationStereo not run
Audio44.1 kHz mono, +288 ms16 kHz stereo, +97 ms

两段都没有接触力、物体 6DoF、机器人末端动作。E6 的 grayscale tracking 与朝下 ctrl 在观测里;ctrl 不是手柄位姿(手柄文件一行零)。E2 没有这些灰度相机。

Neither episode has contact force, object 6DoF, or robot eef action. E6 grayscale tracking and downward ctrl are observations; ctrl is not controller pose. E2 has no grayscale cameras.

07 · 交付样例07 · Delivery samples

原片、对齐、语义与手部估计,按需求加工交付

Source, alignment, semantics and hand estimates — delivered by depth

主推 VLM 包落在 device / human / indexes;estimate 可选。MCAP 与 LeRobot 目录已预留,当前未导出。下面四类是真实交付清单,不是口号。

Primary VLM packs land as device / human / indexes; estimate optional. MCAP and LeRobot folders are reserved, not exported yet. The four cards below are the real delivery checklist.

TOBI DATA PACK 拓比交付包 · 样例字段即交付字段 TOBI pack · sample fields = delivery fields

根据项目选深度。默认 D2;VLM 主推加 D3;手部辅助加 D4。

Pick depth per project. Default D2; add D3 for VLM; add D4 for hand aids.

查看真实交付样例Open real delivery samples

D2 D2 · INDEXES 基础数据 · 标准化采集对齐 Base · standardized aligned capture time_sync、索引、validity。适合先对接管线。 time_sync, indexes, validity. First pipeline hook.
D3 D3 · HUMAN 语义数据 · 场景 / 叙述 / 关键帧 Semantic · scene / narration / keyframes VLM 主推档:人工句子可核对。 Primary VLM tier: checkable human sentences.
D4 D4 · ESTIMATE 行为估计 · 面向策略辅助 Behavior estimate · policy helpers 手点与打码可选;不写成机器人动作或深度真值。 Optional hands / mosaic; not robot actions or depth truth.

01

数据文件

Data files

  • 视频 / 双目或多路原片
  • IMU / 传感器流
  • 时间同步与音频偏移
  • Video / stereo or multi-cam source
  • IMU / sensor streams
  • Time sync and audio offset
MP4CSVJSON

02

结构化结果

Structured results

  • 叙述 / 关键帧 / 物体时间段
  • 手部轨迹(设备或 Vision 估计)
  • 动作时间段(人工核对)
  • Narration / keyframes / object spans
  • Hand tracks (device or Vision estimate)
  • Action spans (human-checked)
JSONLJSON

03

数据索引

Data index

  • Episode / 任务时间戳
  • 场景与文件关联
  • RGB / 手部 / 同步索引
  • Episode / task timestamps
  • Scene and file links
  • RGB / hand / sync indexes
JSONJSONL

04

质量报告

Quality report

  • 数据量与有效率
  • 异常与门禁
  • QA / validity 验收
  • Volume and yield
  • Anomalies and gates
  • QA / validity acceptance
JSONHTML
device/原始流与标定
训练读这里
Raw streams & calibration
Train here
estimate/离线算法
带 method
Offline algorithms
with method
human/叙述与关键帧Narration & keyframes
indexes/时间同步
validity
Time sync
validity
preview/低帧率 mp4
仅供人看
Low-Hz mp4
humans only
用于模型训练Model training · VLM / 场景描述 用于算法评测Evaluation · validity · 工作台核对 用于数据研究Research · 字段与时间轴审计

VLM 最小集:device/ 原片、标定与音频偏移 · human/narration.jsonl(及 E6 关键帧)· indexes/time_sync.json · validity。预览 mp4 与未门控的估计距离不进训练列,公开模型分数也不写入。

Minimum VLM set: device/ source, calibration, and the audio offset · human/narration.jsonl (plus E6 keyframes) · indexes/time_sync.json · validity. Do not put preview mp4, ungated estimated distance, or public model scores into a training column.

episode_bundle.v1.json embodied_layers.v1.json 交付分级说明Delivery tier note

开始构建Get started

三件事,选一件

Three paths — pick one