Data becomes the policy specification

LeRobot is a Python-based OSS that handles multiple robots in a unified manner, and aims to connect datasets, remote control, strategies, training, and visualization. The LeRobotDataset described in the README combines visually synchronized MP4/images with Parquet with state/behavior. The value of having the same format is not just for changing models. The ability to track what command was given to which camera, at what time, and to which joint state.

  1. 1Task Specifications/Safety Conditions
  2. 2Remote Control and Sync Recording
  3. 3Episode Inspection
  1. 1Validated data
  2. 2train/validation/test split
  3. 3imitation learning
  1. 1Policy
  2. 2Simulation evaluation
  3. 3Slow, supervised hardware evaluation
  1. 1Failure Episode
  2. 2Cause Tag
  3. 3Next Collection Design
Consider the sequence and each role.

Things to decide before collecting

The task does not end with just "placing the object," but also defines the starting state, target object, success determination, prohibited contact, maximum speed, stopping action, and shooting position. Observations include not only camera images, but also joint positions and velocities, grippers, task instructions, and time stamps. A time lag between the camera and the action is a typical cause of broken learning, even if it looks good. The image that the operator sees on the collection screen and the saved image are not necessarily the same.

If you list only successful demonstrations, you won't be able to learn how to recover from a slight failure. Explicitly tag failures, interruptions, and anomalies, including different starting positions, lighting, object poses, and late corrections, to the extent safely permitted. Do not intentionally attract dangerous collisions or close contact with people. Before deleting data, record why you excluded it. Being able to modify the label later is one thing, but preserving the original raw data is another.

Expected value of imitation learning

Imitation learning trains a policy to reproduce expert behavior from observations and instructions. When the robot departs from the demonstrator’s trajectory, it can enter states absent from the training data; errors can then compound (exposure bias). Action chunks allow for smoothness and computational efficiency by taking actions several steps at a time, but there is also the risk of continuing old chunks when the environment changes. The frequency of re-observation, mid-course stopping, and speed limits matters on real robots.

Even though LeRobot provides a unified Robot interface, it does not eliminate the calibration, torque limits, range of motion, and communication delays specific to each robot. The list of compatible robots in the README is an entry point for integration and does not imply safety compliance of any configuration. Treat SmolVLA’s published training examples as a reproducible starting point. Before running commands from the official announcement, pin the dataset ID, version, required GPU, license, and evaluation method.

Evaluation and Sim2Real

First, isolate the test split from the time it is collected. When consecutive frames from the same shooting session are divided into train/test, substantial memory is mistaken for generalization. The evaluation includes not only the success rate, but also the starting posture, object, lighting, failure type, number of stops, trial time, and number of interventions. Allow each trial to be recovered with video and telemetry.

Simulations are an inexpensive way to test range of motion, collisions, and task logic. However, friction, stiffness, delay, camera noise, and the way objects appear differ from reality. Domain randomization is a method for expanding differences and learning, but it is not a proof that covers all important differences. Sim2Real is not a one-time port, but a step-by-step hypothesis verification process: simulation → no-power check → low-speed/space isolation → one trial with supervision → log evaluation.

Exercise: Write a data card

Create data cards for hypothetical "color block sorting" without operating the actual machine. Write down the sensor, action unit, frequency, synchronization method, starting state, success, exclusion, privacy, license, and known deficiencies on one sheet. Split 10 hypothetical episodes into train, validation, and test sets, and assign a cause tag to each episode. Next, consider which measurement can detect an abnormality by introducing the camera is delayed by 150ms'', the box is reflective'', and ``the gripper slips once''. I do not claim to have safely operated the actual machine here. This is an exercise in creating records that can be reviewed for design.

Specific example of collection table

Write the logic schema of Parquet/CSV first. episode_id is the immutable ID used for division, frame_index is the order within the same episode, timestamp_ns is the monotonically increasing time, image_path is the visual, state.q is the joint angle, state.gripper is the opening/closing, action.delta_xyz is the hand increment, and action.gripper is the next opening/closing command. Write unit and coordinate_frame together in each column to avoid ambiguity as to whether action.delta_xyz=[0.01,0,0] is 1cm towards +x of the base frame or 1cm towards the front of the tool frame. When re-encoding images or thinning frames, re-examine the correspondence between timestamp and action.

The quality check provides the missing rate, continuous timestamp, state value range, image resolution, episode length, and action saturation rate. Episodes where q is outside the joint limit, action is the same every time, all frames are black, and task is empty are isolated before training and the reason is left. If only the success episodes are extremely short, the model may not have learned when to stop. Whether to remove or include failures depends on the task and safety conditions, but data formatting that does not record what was done is not reproducible.

Train/eval leakage practical check

Save the split manifest first, and fix the seed and version so that the same episode does not move when rerun. train_episode_ids ∩ test_episode_ids = ∅, multiple hashes of the same raw video are not present in the split, and if the same collection session is required, only one side of each group is automatically checked. The average and variance of normalization are calculated only from train, and test statistics are not mixed into preprocessing. When using test for hyperparameter selection, new unseen data is reserved for evaluation.

After evaluation, do not close the failure episode with only success/failure. People add cause tags such as perception, calibration, grasp, trajectory, latency, safety_stop, and unknown after watching the video and telemetry. You can return to the next collection to determine whether the model is bad, data is insufficient, or the success judgment is ambiguous. By leaving the steps to reproduce the failure, the policy and dataset versions, and the environmental conditions on one line, it is difficult to misinterpret a coincidence as an improvement.

Dataset revisions need a replay boundary

A dataset revision is not only a new set of files. Freeze a manifest containing episode IDs, sensor schema, action units, calibration version, and decoder version. Then replay the exact held-out episodes through the loader before training. Count dropped frames, non-monotonic timestamps, missing actions, and normalization values outside the training range. These checks find a broken data path before a loss curve is mistaken for a policy improvement.

The LeRobot repository is a rolling release; a September 2026 merged release relevant to this chapter was not verified. Treat current documentation as an interface to inspect before executing, not as proof that a past command remains reproducible.

Current tooling is not dataset validity

NVIDIA Isaac ROS 5.0, announced September 22, 2026, is a platform announcement about ROS packages and deployment workflows. It does not validate a LeRobot dataset or Sim2Real result. Use the separation operationally: pin the ROS/driver/container versions in the manifest, replay a sensor sample through that exact stack, and fail the run if feature names, units, or timestamps change. Tooling updates should trigger a compatibility test, not a new training claim.

The 2024 data-format baseline

Hugging Face’s August 27, 2024 analysis makes storage and decoding part of the training decision. For this lab, keep the original episode untouched and record a transformed copy’s codec, pixel format, timestamp mapping, and decoder version. Read the same random frame IDs from both representations, then compare observation/action alignment before measuring load time. A smaller file is not a useful improvement if the sample now pairs an image with the wrong action. This is an unexecuted experiment plan, not a reproduction of the publisher’s benchmark.

MENTAL MODEL / COORDINATES

The same point has different coordinates in different frames.

Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.

x = cos θ
y = sin θ

This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.

xy(0.87, 0.50)

SOURCES

01
LeRobot repository ↗github.com · unknown
02
LeRobot documentation ↗huggingface.co · unknown
03
SmolVLA announcement ↗huggingface.co · 2025-06-03
04
NVIDIA Isaac ROS 5.0 announcement ↗blogs.nvidia.com · 2026-09-22
05
Scaling robotics datasets with video encoding ↗huggingface.co · 2024-08-27

YOUR NOTES