Boundaries of this chapter
This article deals with learning procedures that do not connect the robot, do not move the motors, and do not endanger others or property. The safety, electrical wiring, compatibility, and performance of the physical robot hardware have not been verified in this chapter. When starting Physical AI, it is better to be able to read the cause and effect of data and evaluation than to buy a body first, which will make future failures cheaper.
- 1Read public materials
- 2Visualize an episode of data
- 3Languageize success criteria
- 1Build an Evaluator with Synthetic Data
- 2Disprove Hypothesis with Simulation
- 1Safety review
- 2Determine if a real machine is needed
- 3Phase test with another plan if necessary
Week 1: Reading Observations and Actions
Read the public materials of LeRobot and explain in your own words what the images, state, action, episode, and timestamp of the dataset refer to. Draw one public episode in chronological order, and make three columns: changes that can be seen in images,'' changes that can only be seen in state,'' and ``intent expressed by action.'' The first thing to discover here is that you can't tell the gripper opening or joint limits just from the image, and you can't tell what was avoided just from the action.
Next, redefine the success label. Did the object just enter the destination for a moment, was it placed stably, or was the contact within an acceptable range? The success rate of a paper depends on this definition. When reading OpenVLA paper or GR00T N1, make it a habit to extract the evaluation protocol before the model name.
Week 2: Write the evaluator first
Create synthetic 2D block worlds on paper, spreadsheets, or simple programs. The states are block=(x,y), goal=(x,y), obstacle, and the only actions you need to do are move up, down, left and right, and stop. Before learning the strategy, write an evaluator to judge arrival, obstacle collision, boundary deviation, and step count excess. Create three intentionally bad policies and verify that the evaluator detects each failure.
The point of this exercise is not to make the model look good, but to discover first what the evaluation will miss. If the reward is only for being close to the goal, a route that passes through obstacles may result in a high score. If you only look at the exact location, you will miss the dangerous vibrations that occur along the way. Before finding the same hole in the actual machine, you can modify the specifications in a low-risk system.
Third week: Intentionally creating a distribution shift
After limiting the learning background to white and the blocks to red and blue, the evaluation adds yellow, shadows, camera shift, obstacles, and changes to the starting position. When accuracy decreases, don't just say "the model doesn't generalize." Record each change in visual representation, coordinate transformation, behavioral constraints, success definition, and amount of data. Although large-scale pre-training of VLA can aid in semantic representation, it does not obviate the need for sufficient data and evaluation of the physical conditions at hand.
Stopping conditions before proceeding to the actual machine
If the following is unclear, do not start planning for the actual machine. (1) Who can press the emergency stop and when; (2) Upper limits on speed, force, and range of motion; (3) Separation from people and fragile objects; (4) Switching between remote control and autonomous action; (5) Stopping when communication is interrupted; (6) Logging and reproducibility procedures; (7) What is acceptable for failure. High performance model demos do not replace this list. Multi-degree-of-freedom systems such as humanoids require additional risk assessments, including falls and contact.
Summary exercise
Select one published paper and summarize it on one page using six columns: Input'', Output Action'', Body'', Data'', Evaluation'', and Unspecified Safety Boundary''. Next, write three reproduction tasks that do not use the actual machine. Examples are data schema verification, time difference detection between videos and action times, and safety evaluators in synthetic environments. Finally, write the same number of things that cannot be claimed based on this result alone. Reliable progress in learning Physical AI is not the number of demos completed, but the ability to separate observed facts from untested hypotheses.
Four deliverables that can be verified without an actual device
The first is a dataset validator. Read the above schema and set the missing, unknown unit, timestamp inversion, frame that spans episodes, and out-of-range action to report. The second is the replay viewer, which arranges images, states, and actions on the same time axis. It is possible to check not only the visual field at hand, but also ``what commands were given at this time'' at a glance. The third is a policy contract test, which checks whether the policy returns an exception or a safe no-op when given an input shape, an illegal value, and a stop flag. The fourth is a simulator-free counterfactual table that defines the expected transitions when an object disappears, the camera is occluded, or communication is delayed.
These do not prove real-world performance, but reduce ambiguity before proceeding to real-world performance. Leave test_id, input_version, expected, observed, reviewer, timestamp in the verification result. If observations differ from expectations, question whether the specification, data, or evaluator is wrong before correcting the model. Not having an actual machine is a constraint, but it is also an opportunity to learn safety conditions and measurability first.
Numerical Stop Boundary
If action sends delta_xyz at 10Hz, each 1cm period command will be 10cm in 1 second without checking. If we write the upper limit as ||delta|| <= 0.005m and the working area as x∈[0.1,0.4], y∈[-0.2,0.2], z∈[0.05,0.3], we can play it deterministically before passing it to the controller. These values are not recommended values for actual equipment, but are examples to clearly state the boundaries. Actual values are determined by the safety officer based on the model, load, surroundings, and impact on people.
As submissions for the non-actual machine, (a) an example of a schema validator failure report, (b) test results that detected split duplication, (c) round-trip calculation of coordinate transformation, and (d) state transition diagram including stopping conditions. By writing "input", "expectation", "observation", and "unverified" in each deliverable, the next person can review the design without hardware. Treat it as a product that safely reduces unknowns, rather than a product that claims performance.
The review requires one counterexample for each work product. Enter reverse timestamp for validator, duplicate video for split, reverse rotation sign for coordinates, and communication breakdown for state diagram. Confirming only normal cases is not the basis for implementing a dangerous boundary. The records obtained through counterexamples will serve as a basis for deciding whether to start a commercial project in the future.
The conditions for completing the work will also be clearly stated. For each counterexample, save that the test returns a reasoned fail instead of pass, and that the same test returns pass in a version that fixes the failure. Since the results cannot be reproduced by simply checking the display manually, the hash of the input file and a test version are also included. At this stage, we do not evaluate the strategy as ``smart.'' The premise of safe research only confirms what is visible when mechanically broken.
Safety gate: separate a simulator pass from permission to energize
A simulator pass only establishes that a specified observation-action trace did not violate the simulator’s assumptions. Before hardware, write a gate with an owner and observable stop conditions: stale camera frame, missing joint state, action magnitude above a bounded envelope, unexpected contact, lost watchdog, or a human entering the work envelope. The safe action must be explicit—hold, zero velocity, or power-disable—and independently testable without a learned policy.
This is an offline design exercise. A green checklist is not authorization to connect motors, and no acceptance rate turns missing emergency-stop behavior into a software issue.
A recent safety announcement is not a safety case
The September 21, 2026 NVIDIA safety overview argues for lifecycle safety across hardware, software, simulation, and validation. That is an architecture claim, not evidence that a specific lab is safe. Convert it into an artifact: trace each stop condition to a sensor, an independent monitor, a commanded safe state, an owner, and a test record. If one link is absent, the offline exercise stays offline.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
SOURCES
01YOUR NOTES