Time of investigation and how to read it

This is an updated note that confirmed the primary materials released on 2026-10-04. Rather than creating a performance ranking based on the "latest," we leave out what each material actually claims and what it has not yet claimed. The paper's benchmark is the result of data, robots, and success conditions determined by the author, and is not a guarantee for separate bodies, separate rooms, and separate safety conditions.

SmolVLA: Opening an accessible experimental system

SmolVLA paper and Hugging Face presentation present 450M VLAs, public community-derived data, and training and inference procedures through LeRobot. The presentation explained SmolVLM2, flow-matching action expert, and asynchronous inference, and the major point was that it created an entry point with a single consumer GPU and small equipment in mind. The important thing is not to think ``because it's small, it's safe everywhere'', but to be able to observe for yourself the chain of data formats, strategies, and evaluations. Published experimental procedures are the gateway to reproducibility, not success rates with the camera arrangement at hand.

π0 and π0.5: Representing behavior as a flow

π0 superimposes behavior generation using flow matching on the semantic representation of pre-trained VLM, and explores general strategies with data including single-arm, dual-arm, and moving manipulators. The core point is that the continuous action is not fixed to a single discrete token selection, but rather attempts to generate a temporally smooth trajectory. π0.5 considers open-world generalization by co-training different data such as heterogeneous robots, high-level semantic prediction, and the web. "Open-world" here is a research claim in a paper, not a product certification that can be run unsupervised at home. Safety guarantees against unobservable obstacles, fragile objects, and human contact should not be inferred from the title of the paper.

OpenVLA and GR00T N1: Openness and body diversity

OpenVLA is an open source VLA that expands the basis for comparison in which weights, codes, and data conditions can be considered. Check the license, data used, checkpoints, and execution environment separately for each release. GR00T N1 reports a dual-system that connects the visual language module and the behavior module of the diffusion Transformer for humanoid robots, and handles a mixture of real machine trajectories, human videos, and synthetic data. While humanoids may have easier access to tools in the human environment, they also increase the safety issues of freedom of movement, falls, contact, and remote monitoring.

  1. 1OpenVLA
  2. 2Entrance to compare and modify the public infrastructure
  1. 1π0/π0.5
  2. 2VLM Conditional Continuous Behavior and Heterogeneous Data
  1. 1SmolVLA + LeRobot
  2. 2Entrance to test the training and evaluation chain on a small scale
  1. 1GR00T N1
  2. 2Basic Model Study of Cross-Body Including Humanoids
  1. 1Common Unresolved
  2. 2Data Quality/Out of Distribution/Delay/Safety/Reproducibility
Consider the sequence and each role.

Practical comparison questions

Before introduction, make a table that includes (1) Does the target body match the action space? (2) What are the camera, joint states, and frequencies? (3) Which weights, codes, data, and licenses are made public? (4) What is the definition of success and the number of trials? (5) Who will guarantee stopping and resuming in the event of failure. The answer cannot be answered simply by number of parameters,'' demo video,'' and ``first place in benchmarks.''

There are still some unconfirmed matters. This note does not run each model in this environment, and does not verify compatibility with specific hardware, inference delay, license application, or reproducibility performance. When updating, do not change the publication date and add new notes to distinguish between the original claim and subsequent results.

Verification limits for public information in 2026

On 2026-10-04, we investigated each organization's public pages, arXiv, and GitHub. The publication dates that can be confirmed and cited in this note are 2024-06-13 for OpenVLA, 2024-10-31 for π0, 2025-03-18 for GR00T N1, 2025-04-22 for π0.5, and 2025-06-02 for SmolVLA (HF announcement 2025-06-03). In this review, we were unable to add any new foundation VLA announcements for 2026 that meet these same comparison criteria and are supported by primary sources. This is not a conclusion that there will be no announcement in 2026, but rather a limitation of the fact that there was no material that could be confirmed or described within the scope of this investigation.

In the next update, we will not only search for model names but also check each official research blog, repository release, paper revision, and license change separately. Leave source_url, announcement_date, accessed_date, evidence_type, unverified in the evaluation column of the update note. For example, even if an announcement exists, it does not automatically support weight disclosure, reproduction procedures, commercial feasibility, or actual machine testing.

Experiment table to translate comparison into actual work

For a small-scale introduction evaluation, select one model and fix body, camera, state_dim, action_dim, action_horizon, control_hz, dataset_snapshot, success_definition, stop_definition to CSV. Conduct at least multiple trials with different starting positions, and record not only successes but also contact, stoppage, and human intervention line by line. When comparing models, it is impossible to determine which one is better unless they use the same object, the same camera, the same stopping rules, and the same trial budget. At the stage where you do not have an actual machine, fill in only the specified information of published papers in this table, and use the blank spaces as the next research topic.

Compare embodiment shift before score shift

A shared benchmark score does not establish that two policies transfer across bodies. Compare gripper geometry, action horizon, observation stack, teleoperation source, reset protocol, and whether success is measured per trial, per subtask, or after human intervention. The July 2026 community discussion of LingBot-VLA 2.0 reported a large in- versus out-of-distribution gap; that post is a discovery signal, not an independent result. It is useful because it turns “generalist” into a testable split rather than a label.

For a selection sheet, reserve one held-out embodiment or camera placement, freeze the task language, and report success with the number of trials and manual resets. A policy that succeeds after hidden operator recovery belongs in a different column from autonomous completion.

A current announcement changes the question, not the ranking

The September 10, 2026 Skild AI S1 announcement makes a strong in-context-learning claim around long-horizon tasks. Add a column to the 2025 comparison card: what is conditioned at deployment—language, image, video demonstration, retrieved memory, or weights—and which of those inputs was held out in evaluation? A video-conditioned result should not be compared directly with a language-only result until task information and reset budget are matched.

MENTAL MODEL / COORDINATES

The same point has different coordinates in different frames.

Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.

x = cos θ
y = sin θ

This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.

xy(0.87, 0.50)

SOURCES

01
SmolVLA paper ↗arxiv.org · 2025-06-02
02
SmolVLA announcement ↗huggingface.co · 2025-06-03
03
pi0 ↗arxiv.org · 2024-10-31
04
pi0.5 ↗arxiv.org · 2025-04-22
05
GR00T N1 ↗arxiv.org · 2025-03-18
06
Reddit robotics discussion: LingBot-VLA 2.0 ↗www.reddit.com · unknown
07
NVIDIA: Skild AI S1 physical AI announcement ↗blogs.nvidia.com · 2026-09-10

YOUR NOTES