返回数据方案English

SOLUTIONS · E6 + E2

解决方案:把两段样例拆成可训练的数据理解。

交给客户的是标注过的第一人称数据集,用来训练机器人对眼前世界的理解。下面只打开两段已经核对过的记录,不是小时数目录。预览片给人看,训练读原始视频和 jsonl。缺的传感器写成没有,不补一条假轨迹。

E6 · 32 秒办公操作E2 · 11 分 22 秒办公坐姿人的观测,不是机器人动作

先按三件事拆

SCENE

场景

E6 后 10 秒拆成握持手柄、纸卷和布、小零件与看手机。句子是人工关键帧,不是检测器。E2 整段是同一张办公桌,没有逐帧动作名,时间轴是陀螺角速度。

FORMAT

输出格式

mp4 是降帧预览。训练用 jsonl:每一行对齐到源帧,带 dt。手部、文字、物体、陀螺分开存,不混成一种框。

ALGORITHM

算法需求

设备里已经有的,直接交付。设备里没有的,才用后处理,并写上模型名和置信度。E6 的 26 关节是机内 OpenXR。E2 的手是 Vision 估计。

TRAINING

给本体学什么

这两段提供第一人称画面、手和头或惯性怎么一起动。它们不是夹爪 7 维,没有接触力,也没有物体 6DoF,不够单独训一个操作策略。

场景

E6 · 22.00–25.50 s

握持手柄,看向终端

左手一开始还不稳定,之后左右手同时有效。关节来自设备,不是后画的骨架。

打开这段并跳到时间轴

E6 · 25.50–29.50 s

纸卷和布

右手更稳,左手时有时无,然后双手同时有效。物体只有时间段,没有框。

打开 E6 时间轴

E6 · 29.50–32.05 s

小零件,再看手机

先是双手捏持,片尾左手持手机看屏幕。屏幕文字没有逐字识别。

打开 E6 时间轴

E2 · 0–682.17 s

同一张办公桌,11 分 22 秒

左右目 30 Hz。手上的骨架、马赛克和右上视差是后处理。2:17 前后读到墙标「创维XR」和 SKYWORTH。10:28 的蓝色墙标已打码,这一秒 OCR 又是低置信错字。

打开 E2 对比页

同一需求,两台设备给出的理解不一样

算法需求E6 这一包E2 这一包
手部机内 OpenXR 26 关节,约 19 Hz。有可见性,没有逐关节置信度。Vision 21 点估计。不是设备关节。
头怎么动60 Hz 位姿,位置加四元数。5 mm 仍是标称,本段没有地面真值。没有头部位姿。底部角度是陀螺积分,会漂移,不是 SLAM。
看见什么人工四句场景。物体无框、无 mask、无 6DoF。人工每 60 秒抽一帧。没有物体类别名。
文字和人物预览上方做了模糊。不是逐帧人脸检测。人脸、远处人体、置信度达到 0.40 的文字、亮屏幕、蓝色墙标。玻璃淡字未单独打码。
远近没有深度流,不把双目写成深度图。右上是未矫正视差,不是深度相机,也不能当米用。
声音44.1 kHz 单声道。第 0 包比 RGB 第 0 帧晚 288 ms。16 kHz 立体声。播放器按晚 97 ms 叠进标注片。原始对比片无声。

输出文件

用途E6E2
给人看的预览10 Hz,从源 22.0 秒到结束4 Hz 标注片,带声音;旁边是无声原片
训练索引rgb_index.jsonl 1920 行,60 Hzstereo_index.jsonl 20466 行,30 Hz
hands_flags.jsonl写在 e2-findings.jsonl,是估计
场景句narration.jsonl · actions.jsonlnarration.jsonl · motion_spans.jsonl
时钟time_sync.jsontime_sync.json

训练不要读预览 mp4。E6 读 60 Hz 的 rgb.mp4。E2 读两路 30 Hz 原片。手不要重采样成视频帧率而不保留 dt。公开写法上可对照 OpenXR 手、LeRobot / RLDS、Ego-Exo4D、HOT3D、Project Aria。它们是格式对照,不是这两台设备的测量结果。

E6 数据说明 · E2 对比页

SOLUTIONS · E6 + E2

Two episodes, split by scene, format, and algorithm.

What we hand over is an annotated egocentric set for training a robot’s view of the world in front of it. Only two checked recordings are open here. Preview mp4 files are for people. Training reads the source video and the jsonl. Missing sensors stay missing.

E6 · 32 s of desk workE2 · 11 min 22 s seatedHuman observation, not a robot action

Three cuts

SCENE

Scenes

The last 10 seconds of E6 split into gripping controllers, paper and cloth, then a small part and a phone glance. Those sentences are a human keyframe review. E2 stays at one desk. Its timeline is gyroscope rate, not an action name.

FORMAT

Files

The mp4 is a decimated preview. Training jsonl rows join a source frame and keep dt. Hands, text, objects, and gyro stay in separate files.

ALGORITHM

What must be computed

Streams the device already wrote are delivered as-is. Anything added later carries a model name and a confidence. E6’s 26 joints are on-device OpenXR. E2’s hands are a Vision estimate.

TRAINING

What the body can learn

These episodes supply the egocentric image and how the hands move with the head or the inertial sensors. They are not a 7-DoF gripper. There is no contact force and no object 6DoF, so they are not a manipulation policy by themselves.

Scenes

E6 · 22.00–25.50 s

Grip the controllers, face the terminal

The left hand is unstable at first, then both hands are active. Joints come from the device.

Open the E6 timeline

E6 · 25.50–29.50 s

Paper roll and cloth

The right hand is steadier. Objects are time spans, not boxes.

Open the E6 timeline

E6 · 29.50–32.05 s

A small part, then the phone

Both hands hold a part, then the left hand holds a phone. Screen text was not read character by character.

Open the E6 timeline

E2 · 0–682.17 s

One desk for 11 minutes 22 seconds

Stereo at 30 Hz. The skeleton, mosaic, and corner disparity are added later. Around 2:17 the crop reads 创维XR and SKYWORTH. At 10:28 the blue wall mark is masked and OCR is low-confidence again.

Open the E2 comparison

Same need, two devices

NeedThis E6 packThis E2 pack
HandsOn-device OpenXR, 26 joints, about 19 Hz. Visibility yes, per-joint confidence no.Vision 21-point estimate. Not a device joint.
Head motion60 Hz pose. The 5 mm figure is nominal. This clip has no ground truth.No head pose. The bar is integrated gyro and it drifts. Not SLAM.
What is seenFour human sentences. No boxes, masks, or 6DoF.A human look every 60 seconds. No object class names.
Text and peopleThe preview blurs the top band. Not a per-frame face detector.Faces, distant people, text at 0.40 or above, bright screens, and the blue wall mark. Faint glass lettering is not separately masked.
RangeNo depth stream.Unrectified disparity. Not a depth camera and not meters.
Sound44.1 kHz mono. Packet 0 is 288 ms after RGB frame 0.16 kHz stereo. The annotated player delays it 97 ms. The original player is silent.

Files

UseE6E2
Preview10 Hz, from source 22.0 s to the end4 Hz annotated file with sound, beside a silent original
Training indexrgb_index.jsonl, 1920 rows, 60 Hzstereo_index.jsonl, 20466 rows, 30 Hz
Handshands_flags.jsonle2-findings.jsonl, an estimate
Sentencesnarration.jsonl · actions.jsonlnarration.jsonl · motion_spans.jsonl
Clocktime_sync.jsontime_sync.json

Do not train on the preview mp4. E6 trains on the 60 Hz rgb.mp4. E2 trains on the two 30 Hz source files. Do not resample hands up to the video rate without keeping dt. OpenXR hands, LeRobot / RLDS, Ego-Exo4D, HOT3D, and Project Aria are format references, not measurements from these two devices.

E6 note · E2 comparison