What is VLA?

Physical AI does more than interpret camera images and language: it reads the state of a physical body, acts under timing constraints, and receives new observations from the world. VLA (Vision-Language-Action) describes research that joins vision, language, and action in one policy. RT-2 explores transferring web-derived vision-language knowledge to robot control, while OpenVLA makes this class of work available as an open model. A VLA is not a box that looks once and emits the correct motor command. It sits in a closed loop with cameras, coordinate frames, joints, latency, contact, and emergency stops.

  1. 1Human Purpose
  2. 2Language Task Conditions
  3. 3Perception (Image, Depth, State)
  1. 1Perception
  2. 2coordinate frames and world model
  3. 3VLA policy
  4. 4action chunk
  1. 1Action chunk
  2. 2low-level controller
  3. 3robot and physical world
  4. 4next observation
  1. 1Safety monitoring/stop condition
  2. 2Refuse/stop action at any time
Consider the sequence and each role.

Read history in “What did we connect?”

A classical robot estimates the object's pose, converts the hand target into joint angles using inverse kinematics, and moves it using a trajectory tracking controller. While it is explicit and easy to verify, it requires the use of recognizers and rules each time the lighting, objects, or phrases change. Imitation learning reduces this design burden by learning the correspondence between observation and behavior from remote human control. However, it is weak outside the data distribution, and it is impossible to learn how to recover from failure by collecting only successful cases.

Advances in Transformers and vision-language models led to VLA systems that connect language and diverse visual representations to action learning. OpenVLA’s public data and open-source implementation do not by themselves establish that a particular release has available weights, permits commercial use, or is safe on a particular robot. Recent work concerns not only scale, but also how to combine bodies, datasets, and simulation while tolerating latency and real-world failures.

Do not mix the four coordinates

Where beginners stumble is the standard of ``to the right''. Image coordinates are pixels, camera coordinates are three-dimensional from the center of the lens, robot reference coordinates are the position from the base, and hand coordinates are the position seen from the center of the gripper. Calibration defines these transformations. Even though VLA suggests high-level actions, lower-level layers must check reachability, joint range, speed, and collisions. An object that looks like it can be captured in an image may not necessarily be within the arm's range of motion.

What's great and what's unresolved?

By adding language as a condition, it is possible to handle combinations of purposes that have not been programmed individually. By using multiple cameras and state history, it is possible to re-observe partial occlusions and changes in progress. However, this is not proof that we understand the world causally. In some cases, correlations in backgrounds, objects, speeds, and camera positions similar to the training data are picked up. In order to claim safety, it is necessary to measure not only the success rate but also stoppages in the event of failure, contact force, return, unknown objects, lighting differences, and communication delays.

Situations where VLA is used include repetitive tasks involving moving objects, research that requires specifying tasks using natural language, and skill learning from remote control data. It is inappropriate to rely solely on a single end-to-end strategy for motor control that requires millisecond-level determinism, decisions that are directly related to laws and regulations, and situations that require the basis of training data.

Exercise: Disassembly without a body

Without connecting the robot, break down ``Put the red block to the left of the blue box'' on the desk using paper or a diagram. List the necessary observations, each coordinate, success conditions, stopping conditions, and conditions for re-observation after failure. Next, color code the ambiguities that you want to leave to VLA and the constraints that you want to leave in the code/controller. This is an unexecuted design exercise and does not verify the safety of the actual machine. Being able to draw this boundary first is more important than later data collection.

Read observation and action as data

Consider an episode as episode_id, frame_index, timestamp_ns, observation.images.front, observation.state, action, task. observation.images.front is the vision at that time, observation.state is the numerical value such as joint position, speed, gripper state, etc., and action is the value to be commanded in the next control cycle or action chunk. In the format LeRobot describes, visuals can be stored in MP4/images and state/action can be stored in Parquet. The LeRobot README format is a convenient common language, but be sure to write the unit, order, and frequency of each feature on your data card.

For example, even if state=[0.2, -0.1, 0.7, 0.0] is determined as ``x,y,z, gripper'', the meaning will be reversed depending on whether 0.2 is m or a normalized value, whether the origin of z is the desk or base, and whether 0 of the gripper is open or closed. If action=[0.01,0,0,-0.2] is a position increment, it will move forward by 1cm, but if it is a velocity, it will move a different distance after 1 second. Do not pass the paper's action token or normalized value as is to the existing controller.

Check to find leaks

If train/eval is randomly divided by frame, the same object, same background, and continuous hand movements will be included in both, mistaking memory as generalization. Group split at least by episode, and if possible by collection date, object set, and operator. To check, whether the set product of episode_id is empty, whether the perceptual hash of the image file is not duplicated in another split, whether the task statement is too exact a match, and whether the timestamp of test is after train. Once the final tests are visualized and used to fix errors, demote the set to validation instead of locked.

For an exercise without physical hardware, create a synthetic CSV containing 10 episodes and hold out three complete episodes for testing. Deliberately duplicate one frame and put it in train and test to confirm that the hash check stops. Let's first evaluate the ability of data processing to reject bad splits, rather than ``complete the process cleanly.''

Finally, draw the layer that receives the output of the VLA as a state machine. Place the necessary inputs and rejection conditions in each transition: observe → validate → propose → constraint → execute/stop → observe. For example, if the image is missing, it will stop with validate, and if it is out of range of motion, it will stop with constraint. Track this on paper 10 times and check whether there is a transition that "repeat the previous action indefinitely" in any state. This is not a verification of the implemented controller, but an exercise to first clarify the closed-loop safety requirements.

2024–2026 lineage: an action interface is still a contract

The 2024 open-model turn made VLA work inspectable, but it did not remove the interface boundary. For every episode, record camera timestamps, proprioception, action representation (joint target, end-effector delta, or velocity), control frequency, clipping rule, and the frame in which a stop is observed. A model can transfer a semantic instruction while still failing because its action unit or latency differs from the deployment body.

Offline exercise. Take ten public or synthetic episodes and write an action-contract table. Change only one field—camera delay, axis convention, or action scale—in a simulator or spreadsheet. Mark which failures are perception failures and which are contract failures. Do not send commands to hardware.

A September 2026 primary release could not be verified for this article’s cited VLA family as of 2026-10-04; the current boundary is therefore a dated research map, not a claim about the newest deployable policy.

September 2026 signal: in-context task transfer still needs a test split

NVIDIA’s September 10 Skild AI S1 announcement describes a provider claim that a robot foundation model can use a single video demonstration for unfamiliar multistep tasks. It is not an independent benchmark, and it does not make observation-to-action timing portable across bodies. Treat the announcement as a design signal: split demonstrations by object arrangement, camera pose, and task sequence, then measure whether a new demonstration changes action selection without retraining. Record completion, recovery, and unsafe-stop rate separately.

MENTAL MODEL / COORDINATES

The same point has different coordinates in different frames.

Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.

x = cos θ
y = sin θ

This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.

xy(0.87, 0.50)

SOURCES

01
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control ↗arxiv.org · 2023-07-28
02
OpenVLA ↗arxiv.org · 2024-06-13
03
LeRobot ↗github.com · unknown
04
NVIDIA: Skild AI S1 physical AI announcement ↗blogs.nvidia.com · 2026-09-10

YOUR NOTES