Change reading order
When reading a VLA paper, the first thing that catches your eye is the fascinating demonstration and average success rate. However, what is needed for implementation decisions is the contract, not the model name. What is the input, what coordinate system and frequency is the output action, what body and what data are used, what is considered a success, and what is not measured? If you read them in this order, you can see why even though different papers use the same word "VLA", they are not interchangeable.
- 1Thesis Claims
- 2Extract Inputs/Actions/Bodies
- 3Extract Data History
- 1Evaluation conditions
- 2Definition of success and failure
- 3Difference from my conditions
- 1Difference
- 2Reproduction issues/Additional collection/Decision not to introduce
Read π0
π0 connects the pre-trained VLM representation and action generation by flow matching. This study evaluates zero shots, language conditions, and fine-tuning using data including single-arm, dual-arm, and mobile manipulators. The design question here is ``Why do we express action with flow?'' While it has the potential to handle multimodal and continuous trajectories and facilitates smooth motion, it requires the integration of inferential calculations, action chunks, observation delays, and low-level safety constraints. Don't look at the housekeeping demo in the paper and think that you can directly apply the same strategy to defining the joints of another arm or positioning the camera.
Read SmolVLA
SmolVLA highlights small 450M-class VLAs and reproducible experimental routes based on public community data. Publications describe SmolVLM2, multiple RGB, states, language instructions, action experts, and asynchronous reasoning. The big meaning is that ``research on fundamental models is not limited to expensive closed systems.'' Miniaturization may reduce delays, costs, and iterations, but less data does not automatically make it more robust. Public benchmarks and my desk, light, objects, and delays are different distributions.
Read GR00T N1
GR00T N1 combines visual language System 2 and diffusion transformer System 1 for humanoid robots, and mixes real machine trajectory, human video, and synthetic data. The idea of including human videos as a data source is a response to the problem of the scarcity of robot behavior data, but there are still issues such as contact force, joint angles, success or failure of grasping, and how to compensate for physical differences that cannot be obtained directly from video. Being humanoid expands the possibility of adaptation to the human environment, and at the same time expands the verification area for falls, contact, and multiple degrees of freedom.
Comparison yardstick
Do not place the three systems in the same ranking list. π0 focuses on flow action and heterogeneous robot data, SmolVLA focuses on small and public experiment chains, and GR00T focuses on mixing humanoid and heterogeneous data. Fill in the comparison table with body, observation, action representation, data provenance, weights/code/license, training environment, test distribution, number of trials, failure judgment, and actual machine monitoring. A blank field does not indicate a low rating, but rather indicates "I don't know yet."
Exercise: One sheet of introduction evaluation form
Write a judgment on one of the three papers that applies to your hypothetical task Pick a soft bag from the shelf.'' The columns should be observed in the paper'', differences from my environment'', additional conditions for safety'', minimum non-actual machine verification'', and conditions for rejecting adoption''. In the last column, include "There is no data necessary for comparison," "Action definition does not match," and "Cannot implement stop condition." Understanding research is not only a tool for rushing to adopt, but also a tool for creating evidence-based reasons not to adopt.
Template to extract experimental specifications from papers
Immediately after reading, fill in observation={camera names,resolution,history,state}, action={space,unit,horizon,frequency}, dataset={source,episodes,robots,license}, training={pretrain,finetune,compute}, evaluation={tasks,trials,success,intervention}. If the number is not in the text, do not guess and write not reported. Next, categorize the differences from your own design into five columns: input, output, body, distribution, and safety. For example, even if the action in the paper is normalized 7-dimensional, it does not necessarily correspond directly to the tool-frame speed of the 6-axis arm at hand. Check units, order, range, and inversion before writing the adapter.
Minimum unit of comparison experiment
To compare strategies A and B, pair them with the same episode seed, starting pose, object, camera, maximum time, and stopping rule. The result table should be trial_id, policy_version, seed, object_id, start_pose_id, completed, time_s, intervention, stop_reason, video_id. Read not only the difference in success rate, but also the reason for stopping and the frequency of intervention. Determine the criteria for excluding trials in advance, and do not delete videos that fail. If there is insufficient data, do not conclude that there is a difference, but record the insufficient number of sections or trials.
train/eval readability check for leaks
When reading papers and repositories, check whether the same object, the same scene, or the same remote control session spans train and eval, and what unknowns are defined as. A split that changes only the language does not measure visual generalization, and a split that changes only the background does not measure operational generalization. In research using the web or human videos, check the possibility that data similar to downstream evaluations will be included in prior learning, within the scope of what is written in the materials. If it is not written down, it will not be concluded that it is a leak, and it will be held as a limit for comparison.
For verification without an actual device, use the published task definition, group split the composite episode, and intentionally mix the same video hash into train/test. If a split test reports a failure, it at least prevents duplication from being visible to the evaluation pipeline. Although it will not replicate the capabilities of VLA, it will allow you to see if you have been able to translate the evaluation contract of a paper into your own work.
The final result of reading comprehension is not a declaration that ``I reproduced it'' but an implementation work plan. Separately write the environment, data source, license confirmation, required sensors, action adapters, evaluation metrics, stop rules, and information that cannot be disclosed. If either is missing, do not start the replication process and return the missing part to the research question. Not reprinting research numbers as your own is also a basic quality control that connects Physical AI to practice.
At review meetings, only work sheets are handed out to people who have not read the paper. If you cannot answer the following questions: What should be downloaded?'' What should not be downloaded yet?'' How should one trial be considered a success?'' What should be saved when stopped?'' If you cannot answer the questions, revise the work form. An explanation of the novelty of the model and an explanation of how to safely operate the experiment should be separate texts.
Add an embodiment row to every paper card
For each paper, add one row that cannot be averaged away: body, end effector, camera placement, action parameterization, reset procedure, and evaluator intervention. A score without this row can compare semantic competence while hiding a different control problem. Ask whether a failure is counted after the first unsafe trajectory, after a reset, or only after a human has restored the scene.
The reading output should be a falsifiable transfer hypothesis: “under this camera delay and this action scaling, the policy will preserve object orientation for N trials.” If the paper does not publish the needed variables, mark the transfer claim untested rather than estimating them from demo video.
Read deployment announcements as interfaces, not evidence tables
NVIDIA Isaac ROS 5.0’s September 22, 2026 announcement is useful to identify an integration layer between policies and ROS deployment. It should not be placed in a performance table beside π0, SmolVLA, or GR00T. Add a separate “integration evidence” field to the reading card: runtime, message schema, hardware support, version pin, and whether a claimed policy evaluation actually used that path.
Keep the 2024 model claim separate from the platform claim
The June 13, 2024 OpenVLA preprint provides a dated model-level baseline. The September 2026 Isaac ROS announcement above describes a different layer of the system. A platform integration does not reproduce a policy result. Make a two-column evidence sheet: checkpoint, training data and action representation on one side; middleware revision, timestamps and deployment interfaces on the other. For every claimed improvement, identify which column changed and which outcome was measured. If the source never measures the combined stack, mark that comparison unverified rather than combining two separate announcements into one result.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
SOURCES
01YOUR NOTES