Dataset format
The LeRobotDataset layout the hub reads, which versions the episode viewer supports, and what it does with the rest.
A dataset repo on MouseMouse is just files. Push whatever you like — but if the layout is a LeRobotDataset, the hub recognises it and the repo page grows an Episodes tab: the recorded video and the robot posed in MuJoCo side by side, per-joint charts grouped by limb underneath, all on one clock; click a chart to seek.
The Sim replay pane needs to know which robot recorded the data. It reads
info.json's robot_type and looks for a robot repo with that slug under the
same owner (then any public robot with that slug); that robot's card supplies
the MJCF, and its embodiment interface supplies the joint order the
observation.state columns map to (by name, or by position when the columns
are unnamed and the counts match) and the parts[] the charts group by. When
any link in that chain is missing the pane says which one, and the charts fall
back to grouping by feature.
We do not define a format. LeRobot's is the upstream spec and its docs are the reference; this page only says what the hub does with it.
The layout the hub looks for
meta/info.json the manifest — fps, features, and the path templates
meta/episodes.jsonl one line per episode: index, length, tasks
data/chunk-000/episode_000000.parquet
videos/chunk-000/observation.images.cam/episode_000000.mp4
Nothing is version-sniffed. The hub reads info.json's own data_path
and video_path templates, formats them, and checks the result against the
files actually in the repo. A layout that follows its own manifest works, even
if it is not exactly any released version; a layout that contradicts its
manifest is reported as such rather than half-rendered.
Two families are supported:
- Per-episode (v2.0 / v2.1) — one parquet and one mp4 per episode.
- Shared (a v1 monolith, or a v3 data file) — one parquet holds many
episodes and the rows are found by the
episode_indexcolumn. Parquet row-group statistics are used to skip groups that cannot contain the episode, so reading one episode out of a fifty-episode file does not read the file.
Measured on a 50-episode, 15,254,503-byte shared parquet: one episode costs 505,164 bytes — 3.3% of the file — in six ranged GETs. Nothing downloads a dataset to draw a chart.
What is not supported
v3 shared video files — one mp4 holding several episodes back to back. The
hub names them in a warning on the Episodes tab instead of guessing: the frozen
FrameFile schema has no time-offset field, so a shared video could only be
mapped by inventing one, and a viewer that silently shows you the wrong
episode's footage is worse than one that says it cannot.
Everything else degrades to a stated reason. A repo whose files are not a
readable dataset answers 409 with what was wrong (meta/info.json missing,
templates that resolve to files that are not there, an empty episode index),
and the tab shows that sentence rather than an empty state or a crash.
What the charts actually show
The API returns a flat {t, name, value} series for one episode, grouped into
observation.state and action. Two decisions worth knowing:
- Downsampling is min/max decimation, not a stride. A stride drops a
one-frame torque spike — which is the thing you opened an episode viewer to
find. Every point returned is a real sample;
downsampleandsampled_pointsin the response say exactly what happened. - A read budget is enforced per request (
DATASET_READ_BUDGET_BYTES, 64 MiB by default). A pathological file answers409, never an OOM.
Channel colour is assigned per group, so knee in the positions chart and
knee in the actions chart below it are the same hue. That pairing is the
whole reason the two charts are stacked.
Pushing one
lucen push you/my-dataset ./my-dataset --kind dataset --public
Order does not matter and neither does completeness — the tab appears as soon as the manifest and the files it names are both present. Identical bytes are stored once across every repo on the hub, so pushing a dataset you pulled from someone else moves no bytes at all.
Dataset-level metadata (episode_count, total_frames, fps) is a separate
typed write — PUT /v1/datasets/{owner}/{repo}/meta, see the
API reference. It is what the repo cards in listings show, and it
is independent of the files, so a dataset can advertise its size before every
shard has landed.
Importing one from Hugging Face
A LeRobot dataset that already lives on the Hugging Face Hub comes across in
one command. It reads the dataset's meta/info.json on the Hub first and
refuses, with the reason, anything the hub would not read (no info.json, a
codebase_version other than v2.0, v2.1 or v3.0); prints the plan; fetches the
files to your machine through the Hub's own client and pushes them with the
same content addressing as above, so a second run moves nothing; writes the
dataset meta from info.json; and has the hub read the episodes back before
the repo can be made public:
lucen datasets import hf://lerobot/pusht
The repo's description records where it came from, the revision and the
licence the dataset card declares; the README.md is the dataset's own card
when it has one. From a terminal
shows the plan it prints and the question it asks.
Upstream
- LeRobot documentation — the format, the recorders, the policies.
- LeRobotDataset v3 — the current layout and what changed from v2.1.