What this workshop produces
The goal is not an impressive robot video. It is a reproducible evaluation packet: a split manifest for public or synthetic episodes, a fixed baseline policy, one narrowly stated adaptation hypothesis, a failure taxonomy, and an evidence-based Sim2Real readiness decision. This chapter does not connect hardware or run training. Commands and checkpoint versions change, so check the current LeRobot documentation and the policy provider’s primary documentation before executing a run.
- 1Task contract
- 2episode inventory
- 3group split and leakage checks
- 1Fixed baseline evaluation
- 2adaptation hypothesis such as LoRA
- 3evaluate under identical conditions
- 1Failure video and telemetry
- 2causal categories
- 3Sim2 Real readiness decision
- 1Not ready
- 2improve data, evaluation, or safety boundary
- 3evaluate again
0. Freeze a task contract first
For a task such as putting a red cube in a container, document the initial cube, container, and arm pose; the front and wrist cameras; joint and gripper state; and the action representation. Define success as a cube that is stationary in the container, a released gripper, and completion before a time limit. Define failure as a drop, leaving the permitted region, a collision, timeout, or stop. Whether action.delta_xyz is meters per step or meters per second, and whether a coordinate is expressed in the base or tool frame, belongs to the contract too.
Without this document, an apparent improvement cannot be measured. Narrowing start poses can improve success rate, and removing failures from a log can make a system look safer. The contract is both an experiment manual and a boundary against changing the comparison after seeing a result.
1. Inventory episodes and split by groups
For every episode collect episode_id, session_id, operator, object set, camera setup, task text, starting pose, raw-video hash, start time, length, and outcome. Do not randomly split frames. Split by session_id or episode_id, because lighting, background, objects, and operator habits within one session are correlated and create an illusion of generalization when they appear in more than one split.
Run four mechanical checks. Episode IDs and raw-video hashes must not intersect between splits. Object sets or start poses reserved for test must not accidentally return to training. Finally, fit normalization statistics, image-augmentation randomness, and tokenizer vocabulary on training data alone. Computing a statistic over the whole test distribution reveals evaluation information before inference even when no model weights were inspected.
2. Keep a baseline policy fixed
A baseline is the comparison before adaptation; it is not something to remove because it performs poorly. Pick one public policy or simple behavior-cloning implementation and record checkpoint ID, code commit, dataset snapshot, input transform, control frequency, and seed. A paper’s reported result, including one for SmolVLA, is not a result for this task contract.
For every trial retain trial_id, policy and dataset versions, split, object and start-pose IDs, seed, completion flag, elapsed time, stop reason, intervention, and video ID. Where possible, label success and failure without showing the evaluator the policy name. Before optimizing, count the three largest failure categories: an unexpected weak baseline can be caused by camera cropping, action units, temporal alignment, or the start state rather than capacity.
3. Make lightweight adaptation a testable hypothesis
Parameter-efficient tuning such as LoRA learns low-rank updates rather than changing all large-model weights. That does not mean every VLA checkpoint supports it, or that small data produces a safe improvement. Confirm official implementation support, license fit, the target layers, and how the action head is handled before calling it an option.
State one hypothesis, for example adaptation to lighting variation. Compare baseline and adapted policy with the same train groups, validation partition, augmentation, and trial budget. Select rank, learning rate, and steps on validation; touch a locked test set once at the end. An improvement on backgrounds close to training is not generalization. Increased stops, oscillatory actions, and regressions deserve the same prominence as a success-rate gain.
4. Structure failure evidence
Use categories such as perception, localization, grasp, planning, control, recovery, safety_stop, and unknown. A trial can have several causes, so keep a primary and secondary cause with the supporting video timestamp and telemetry. Do not force unknown cases into a familiar category: unknown is a reason to collect a better observation.
Build an annotation confusion matrix. Two reviewers can independently label the same video and reconcile disagreements. Preserve definitions, labeling instructions, and the reason for reconciliation. If reviewers know a model identity, expectation can influence the selected cause; a saved process lets later work test an impression against evidence.
5. Treat Sim2Real as a gate
Simulation success is not evidence of hardware success. Friction, grasping, flexible objects, lighting, camera latency, encoder error, and communication loss are difficult for a simple simulator to represent. Check a set of gates instead: the task contract, action units, coordinate transforms, stop states, data licenses, leakage checks, structured failures, and reproducible evaluation logs. If any is absent, the result is not ready and no hardware test should begin.
A hardware organization also needs an approved safety plan covering emergency stop, velocity and force limits, exclusion zones, supervision, communication-loss behavior, and post-trial inspection. This chapter grants none of that approval. It narrows the unknowns with public or synthetic data so a future safety review has a concrete, reviewable starting point.
Submission checklist
Place the split manifest, leakage-check output, identically formatted baseline and adapted-policy tables, failure counts, open questions, and Sim2Real decision in one folder. Include at least one reason a favorable result is still not ready, such as no independent camera-calibration check, test data from one collection day, or missing stop logs. That honest gap defines the next data collection and the safe hardware plan.
A readiness decision needs a counterfactual log
For every rejected rollout, retain the observation window before the failure, the proposed action, the safety filter’s decision, and the reset required to continue. This lets a reviewer distinguish “the policy chose badly” from “the camera was stale” or “a limit stopped a dangerous action.” Report intervention rate and time-to-safe-state beside task success; otherwise a high success score can hide repeated near misses.
The exercise remains offline until the stop path, watchdog, action envelope, and recovery owner have been tested independently. Sim2Real readiness is a decision record, not a threshold emitted by a benchmark.
Add a safety-case column to the packet
The September 21, 2026 NVIDIA safety overview supports a useful distinction: a policy score is not a deployment safety case. Add one column per test condition for monitor coverage, stop latency, safe-state evidence, and unresolved hazards. A passing score with missing monitor coverage remains “not ready,” even if the task video looks stable.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
SOURCES
01YOUR NOTES