Skip to content
Docs menu

Dataset format

The LeRobotDataset layout the hub reads, which versions the episode viewer supports, and what it does with the rest.

A dataset repo on MouseMouse is just files. Push whatever you like — but if the layout is a LeRobotDataset, the hub recognises it and the repo page grows an Episodes tab: the recorded video and the robot posed in MuJoCo side by side, per-joint charts grouped by limb underneath, all on one clock; click a chart to seek.

The Sim replay pane needs to know which robot recorded the data. It reads info.json's robot_type and looks for a robot repo with that slug under the same owner (then any public robot with that slug); that robot's card supplies the MJCF, and its embodiment interface supplies the joint order the observation.state columns map to (by name, or by position when the columns are unnamed and the counts match) and the parts[] the charts group by. When any link in that chain is missing the pane says which one, and the charts fall back to grouping by feature.

We do not define a format. LeRobot's is the upstream spec and its docs are the reference; this page only says what the hub does with it.

The layout the hub looks for

meta/info.json          the manifest — fps, features, and the path templates
meta/episodes.jsonl     one line per episode: index, length, tasks
data/chunk-000/episode_000000.parquet
videos/chunk-000/observation.images.cam/episode_000000.mp4

Nothing is version-sniffed. The hub reads info.json's own data_path and video_path templates, formats them, and checks the result against the files actually in the repo. A layout that follows its own manifest works, even if it is not exactly any released version; a layout that contradicts its manifest is reported as such rather than half-rendered.

Two families are supported:

  • Per-episode (v2.0 / v2.1) — one parquet and one mp4 per episode.
  • Shared (a v1 monolith, or a v3 data file) — one parquet holds many episodes and the rows are found by the episode_index column. Parquet row-group statistics are used to skip groups that cannot contain the episode, so reading one episode out of a fifty-episode file does not read the file.

Measured on a 50-episode, 15,254,503-byte shared parquet: one episode costs 505,164 bytes — 3.3% of the file — in six ranged GETs. Nothing downloads a dataset to draw a chart.

What is not supported

v3 shared video files — one mp4 holding several episodes back to back. The hub names them in a warning on the Episodes tab instead of guessing: the frozen FrameFile schema has no time-offset field, so a shared video could only be mapped by inventing one, and a viewer that silently shows you the wrong episode's footage is worse than one that says it cannot.

Everything else degrades to a stated reason. A repo whose files are not a readable dataset answers 409 with what was wrong (meta/info.json missing, templates that resolve to files that are not there, an empty episode index), and the tab shows that sentence rather than an empty state or a crash.

What the charts actually show

The API returns a flat {t, name, value} series for one episode, grouped into observation.state and action. Two decisions worth knowing:

  • Downsampling is min/max decimation, not a stride. A stride drops a one-frame torque spike — which is the thing you opened an episode viewer to find. Every point returned is a real sample; downsample and sampled_points in the response say exactly what happened.
  • A read budget is enforced per request (DATASET_READ_BUDGET_BYTES, 64 MiB by default). A pathological file answers 409, never an OOM.

Channel colour is assigned per group, so knee in the positions chart and knee in the actions chart below it are the same hue. That pairing is the whole reason the two charts are stacked.

Pushing one

lucen push you/my-dataset ./my-dataset --kind dataset --public

Order does not matter and neither does completeness — the tab appears as soon as the manifest and the files it names are both present. Identical bytes are stored once across every repo on the hub, so pushing a dataset you pulled from someone else moves no bytes at all.

Dataset-level metadata (episode_count, total_frames, fps) is a separate typed write — PUT /v1/datasets/{owner}/{repo}/meta, see the API reference. It is what the repo cards in listings show, and it is independent of the files, so a dataset can advertise its size before every shard has landed.

Importing one from Hugging Face

A LeRobot dataset that already lives on the Hugging Face Hub comes across in one command. It reads the dataset's meta/info.json on the Hub first and refuses, with the reason, anything the hub would not read (no info.json, a codebase_version other than v2.0, v2.1 or v3.0); prints the plan; fetches the files to your machine through the Hub's own client and pushes them with the same content addressing as above, so a second run moves nothing; writes the dataset meta from info.json; and has the hub read the episodes back before the repo can be made public:

lucen datasets import hf://lerobot/pusht

The repo's description records where it came from, the revision and the licence the dataset card declares; the README.md is the dataset's own card when it has one. From a terminal shows the plan it prints and the question it asks.

Upstream