STEREO
左右目
Camera0 左、Camera1 右。HEVC 1600×1200。标定 group 写的是 tracking。左右曝光起点差的 p50 是 198 微秒,曝光中点差的 p50 是 210 微秒。曝光时长约 4–10 ms,增益固定 1024。针孔估算水平视场约 117°,没用畸变系数,不是产品标称。
E2 · EPISODE 20260917_133640
设备 UID d36997d92fc9,序列号 126091102913。播放器上的骨架、马赛克和右上色块是模型估计,不是这台设备输出的手部、位姿或深度。设备里实际有的仍是左右目、IMU、磁力计和音频。这一包没有头部位姿和 SLAM,画面底部的航向是陀螺积分,会漂移。
E2 · EPISODE 20260917_133640
Device UID d36997d92fc9, serial 126091102913. The skeleton, mosaic and corner patch are model estimates, not device hands, pose or depth. The device files are still the two eyes, IMU, magnetometer and audio. This package has no head pose and no SLAM. The heading in the bar is integrated gyro and it drifts.
正在缓冲这一段
两路都是设备 16 kHz 立体声,比左目曝光起点晚 97 毫秒。一起播放时只响你按下的那一路,避免叠在一起。手机上先看到标注片。马赛克盖人脸、远处人体、置信度达到 0.40 的文字、亮屏幕,以及蓝色墙标。
源 0:00 · 0.00 s · 第 0 帧 · 30 Hz
还没有读到这一秒的记录。
No record for this second yet.
墙上那块蓝色标志按颜色连通域打码:宽至少 160 像素,蓝像素占比约 0.15 到 0.55,位于画面上半部。10:28 左右目各有一块。整段记录里这样的块出现了 32 次,不只这一秒。这不是手工涂在单帧上。
OCR 分成两档。达到 0.40 才采用。整段反复读对的是 SKYWORTH。2:17 前后,标志裁切上还读到「创维XR」。10:28 颜色打码仍在,但这一秒 OCR 又落回低置信错字,例如 SKYWORH。其余达到 0.40 的多是碎字符,不当成物体名。玻璃上很淡的字没有单独打码。
手以播放器里的 Vision 骨架为准,不是设备关节。人脸和远处人体只在检出时打码。纸箱、门、鞋、天花板没有类别框。陀螺时间轴只表示头戴角速度,不是动作名。
The blue wall mark is masked by a color region: at least 160 pixels wide, blue-pixel share about 0.15 to 0.55, in the upper half of the frame. At 10:28 both eyes have one region. The log has 32 such rows, not only that second. It is not a hand-painted patch.
OCR is accepted only at confidence 0.40 or above. The string that repeats is SKYWORTH. Around 2:17 the crop of the mark also reads 创维XR. At 10:28 the color mask is still on, but OCR falls back to low-confidence errors such as SKYWORH. Other accepted strings are fragments, not object names. Faint glass lettering is not separately masked.
Hands in the player are a Vision skeleton, not device joints. Faces and distant people are masked only when detected. Boxes, the door, shoes and the ceiling have no class boxes. The gyro timeline is head-worn angular rate, not an action name.
时间轴仍是陀螺模长,不是动作名。上面“这一秒”来自模型记录:手、蓝色墙标、采用的文字、低置信 OCR、人脸、人体和亮屏。纸箱和门没有类别框。玻璃上很淡的字这次没有单独打码。
The timeline is still gyroscope magnitude, not an action name. The line above comes from the model log: hands, blue wall marks, accepted text, low-confidence OCR, faces, people and bright screens. Boxes and the door have no class labels. Faint glass lettering is not separately masked.
time_sync.jsonstereo_index.jsonlmotion_spans.jsonlnarration.jsonle2-findings.jsonl
STEREO
Camera0 左、Camera1 右。HEVC 1600×1200。标定 group 写的是 tracking。左右曝光起点差的 p50 是 198 微秒,曝光中点差的 p50 是 210 微秒。曝光时长约 4–10 ms,增益固定 1024。针孔估算水平视场约 117°,没用畸变系数,不是产品标称。
IMU + MAG
加计和陀螺约 798.9 Hz,加计模长抽查均值 9.83 m/s²。磁力计正好 200 Hz。加计比左目第 0 帧曝光起点早 124 ms。标定里的 1.5 ms 相机对齐没有套用,因为这包没有头部位姿可以核对。
AUDIO
PCM 16 kHz 立体声。第 0 包比左目曝光起点晚 96.6 ms。不要把 audio.mp4 和视频从 0 秒硬叠。
NOT IN THIS PACKAGE
设备文件里没有手部关节、头部位姿、SLAM、深度和物体类别。播放器里的手是 Vision 估计,右上是视差不是深度相机,底部航向不是 SLAM。
A · DEVICE · CAN TRAIN
训练读 Camera0 与 Camera1 的 30 Hz 原片,共 20466 帧,间隔大于 40 ms 的缺口是 0。不要用这支 4 Hz 预览训练。这是人的第一人称观测,not_robot_action 为真,不能当成机器人末端动作。
Train on the 30 Hz Camera0 and Camera1 files, 20466 frames, with no gap longer than 40 ms. Do not train on this 4 Hz preview. This is a human egocentric observation. not_robot_action is true. It is not a robot end-effector action.
基线 65.94 mm 来自外参位置差,不是棋盘。针孔水平视场约 117°,没用畸变。在补标定之前只做相对立体,不做米制深度。
The 65.94 mm baseline is an extrinsics position difference, not a checkerboard. Pinhole horizontal field is about 117°. Distortion was not used. Until a new calibration, use relative stereo only, not metric depth.
加计比左目曝光起点早 124.3 ms。陀螺 5 秒窗 RMS 中位 0.352 rad/s。磁力计 200 Hz,室内磁干扰没有另测。工厂对齐常数在 time_sync.json 里,本页没有套用,因为没有头部位姿可核对。
Accel leads the left-eye exposure start by 124.3 ms. Median 5 s gyro RMS is 0.352 rad/s. Magnetometer is 200 Hz. Indoor magnetic disturbance was not measured. Factory alignment constants are in time_sync.json and were not applied here: there is no head pose to check them.
立体声第 0 包晚 96.6 ms。可以做键击、说话这类弱对齐,不能把音频和视频从 0 秒硬叠。
Stereo packet 0 is 96.6 ms late. It can weakly align key clicks or speech. Do not stack audio and video at t=0.
左右目硬同步、曝光起点差 p50 198 µs,适合做表征预训练。手、屏、人的框都不是这层的标签。
The eyes are hard-synced. Exposure-start difference is 198 µs at the median. That supports representation pretraining. Hands, screens and people are not labels of this layer.
B · MODEL ESTIMATE · NOT GROUND TRUTH
来自标注片的 Vision 记录,一秒一行。手是 21 点估计,不是设备关节。亮屏是亮度格子,不是显示器检测器。这些数不能当地面真值。
These rows come from the annotated preview, one per second. Hands are a 21-point estimate, not device joints. Bright screens are a luminance grid, not a monitor detector. Do not treat the counts as ground truth.
蓝色墙标 32 秒、合计 57 块。采用文字里 SKYWORTH 出现 8 次。记录的 weak 字段里,「创维XR」有 12 次置信度达到 0.40,页面上的「这一秒」会当成采用。其余达到 0.40 的多是碎字符,不是物体名。纸箱、门、键鼠没有类别框。
Blue wall marks: 32 seconds, 57 regions. SKYWORTH appears 8 times in the accepted-text field. In the weak field, 创维XR reaches confidence 0.40 twelve times, and the per-second line treats that as accepted. Other strings at 0.40 are fragments, not object names. Boxes, the door, the keyboard and the mouse have no class boxes.
下面四段是人工看过标注预览,或只看陀螺。点一下跳到该秒。不是检测器,也不是全长动作名。
The four windows below are a human look at the annotated preview, or gyro only. A click seeks there. They are not a detector and not action names for the whole take.
C · STILL MISSING FOR A ROBOT POLICY
下面的论文是别人对数据字段的要求,不是本段的实验结果,也不是本机测出来的精度。
The papers below are other groups' field requirements. They are not results of this episode and not a measured accuracy of this device.
Open X-Embodiment / RT-X(arXiv 2310.08864)把观测和机器人动作绑在一起。本段只有人的观测,没有末端动作、接触力或语言指令。
Open X-Embodiment / RT-X (arXiv 2310.08864) pairs observations with robot actions. This take has a human observation only. No end-effector action, contact force, or language instruction.
HOT3D(CVPR 2025,arXiv 2406.09598)要的是三维手和物体。本段 Vision 21 点没有逐关节置信度,也没有手套或 MANO。
HOT3D (CVPR 2025, arXiv 2406.09598) asks for 3D hands and objects. The Vision 21 points here have no per-joint confidence, and there is no glove or MANO.
Ego-Exo4D(CVPR 2024)用密集旁白。本段 narration.jsonl 是大约每 60 秒抽看一次,方法写的是人工抽样,不是模型输出。
Ego-Exo4D (CVPR 2024) uses dense narration. narration.jsonl here is a human sample about every 60 seconds, not a model output.
导航和多步操作要头部位姿或 VIO。本段没有。底部航向是去偏后的陀螺积分,轴没有和相机核对,会漂移。
Navigation and multi-step manipulation need head pose or VIO. This take has neither. The heading is bias-subtracted integrated gyro. Axes were not checked against the camera, and it drifts.
右上色块是未矫正鱼眼上的粗 SAD,没有用焦距和基线换成米。抓取位姿需要另做棋盘或深度相机。
The corner patch is coarse SAD on unrectified fisheye. Focal length and baseline were not turned into meters. A grasp pose needs a new checkerboard or a depth camera.
标注片盖了检出的人脸、远处人体、0.40 以上的文字、亮屏和蓝色墙标。77 秒人脸、221 秒人体、642 秒亮屏说明办公画面里这些很多。玻璃上很淡的字可能还在。这不是可直接交易的干净流。
The annotated preview masks detected faces, distant people, text at 0.40 or above, bright screens and the blue wall mark. 77 face seconds, 221 person seconds and 642 bright-screen seconds show how often they appear. Faint glass lettering can remain. This is not a clean file ready to sell as fully masked.
FOUR USERS · ONE EPISODE
同一段 20260917_133640。差别只在你要拿它训练什么。样例卡里的数字都来自本段文件,不是产品标称。
The same episode, 20260917_133640. The difference is what you train. Numbers on the cards come from this episode's files, not from a brochure.
给做自监督表征、双目和时间对齐的人。这一段现在就能进加载器。它是人的观测,不是机器人动作。
For people training self-supervised video, stereo, and time alignment. This take can enter a loader now. It is a human observation, not a robot action.
给以后要做技能时间轴的人。下面是候选窗口,不是已经标好的动作名,也不是逐帧检测。
For people who will build a skill timeline later. The windows below are candidates, not finished action names and not a per-frame detector.
给要做立体匹配的人。这一段能交相对几何。米制深度还没有,因为基线不是棋盘测的。
For people running stereo matching. This take can hand over relative geometry. Metric depth is not here, because the baseline was not measured on a checkerboard.
给要确认能不能下单、字段写不写得清的人。这一页就是数据卡。标注片上的马赛克不是一条已经遮全的干净流。
For people checking whether the fields are clear enough to order. This page is the data card. The mosaic on the annotated preview is not a fully cleaned stream.
四种用户都遵守同一条:只读 Camera0 / Camera1;IMU 早 124.3 ms、音频晚 96.6 ms;骨架和 OCR 标成模型估计;不要把本段当机器人动作监督。
All four users share one rule: read only Camera0 and Camera1; IMU is 124.3 ms early and audio is 96.6 ms late; mark skeletons and OCR as model estimates; do not supervise a robot action with this take.