What the new paper actually measures

The first version of Measuring the Stability Assumption Behind Action Chunking was submitted to arXiv on October 1, 2026. The authors ask a narrow but useful question: after a small injected error in an action, does the state difference shrink or grow? They compare two counterfactual trajectories. In the open-loop branch, the remaining actions from a recorded chunk are replayed without replanning. In the closed-loop branch, the policy receives the perturbed state and replans on later steps. Their rate is fitted over a finite observation window, with states classified as stable, unstable, or unresolved when the evidence is not decisive.

Those are the authors’ measurements in their stated simulation benchmarks and policy conditions. They do not establish a best chunk size for every robot, show that a short chunk is always safer, or report a result for this curriculum. The paper itself treats the fitted rate as a finite-horizon summary and separates error propagation from task performance.

This distinction matters because “the policy emitted eight actions” mixes two questions: how much a first action can move the task away from its nominal path, and whether later observations can recover from that move. A smooth demonstration can conceal either failure. A replayed remainder may preserve an obsolete plan; a replanning policy may still fail if the observation, task state, or recovery data are insufficient.

An unexecuted offline fixture

The following is original educational design inspired by the paper’s measurement framing. It has not been run here. It uses only a simulator or already-recorded, non-actuating rollouts. It does not connect a robot, command motors, train a model, or grant permission for a hardware test.

  1. 1Offline task contract and action envelope
  2. 2nominal trajectory and action stamp
  3. 3bounded synthetic perturbation
  1. 1Nominal branch
  2. 2replay remaining stored actions
  3. 3open-loop deviation
  1. 1Perturbed state
  2. 2replan only in an offline simulator
  3. 3closed-loop deviation
  1. 1Paired traces
  2. 2evidence and permission review
  3. 3next offline question or stop
Consider the sequence and each role.

First write the task contract: state representation, action units and frame, action envelope, success condition, distance measure, and a named reviewer. Pick a stored trajectory and an action stamp. Make a nominal branch with the unmodified action. In a separate synthetic branch, change only one selected action by a predeclared bounded perturbation. The bound must be expressed in the same units and frame as the action; an unexplained “small” displacement is not reproducible.

For the open-loop comparison, replay the remaining stored actions in the simulator without calling the policy. For a closed-loop comparison, ask the policy for later actions only if a resettable simulator can supply the perturbed observations. If the data contain no valid simulator state, no observation model, or no policy access, record closed_loop_unavailable. Do not substitute a guess that replanning would recover.

Before looking at outcomes, choose a distance such as object-pose error, end-effector error, or a task-state distance, and choose a fixed horizon. Save the nominal and perturbed time series, the horizon, perturbation vector, simulator and policy versions, random seed, and any reset. Label the result contracting, expanding, or unresolved only under thresholds stated in the task contract. The label describes this trace under these conditions; it is not a general stability certificate.

Permission and stop are part of the result

The fixture should log four decisions beside its error curve:

Question Required record Safe outcome when evidence is absent
Is the perturbation inside the offline action envelope? units, frame, bound, reviewer stop the case
Can the branch be reset and reproduced? state snapshot, seed, simulator version mark unresolved
May a closed-loop branch call the policy? explicit offline permission and input contract do not replan
Does any proposed follow-up leave the task boundary? stop reason and owner stop; no hardware escalation

An expanding open-loop trace can motivate a review of observation timing or chunk horizon. It cannot automatically shorten a controller’s chunk, authorize a live trial, or prove that the closed-loop policy is unsafe. A contracting trace is equally limited: repeated perturbations, sensor delay, contact changes, and a different task may reverse the conclusion. Keep rejected and stopped cases; excluding them makes recovery look better than the available evidence supports.

How this extends the existing chapter

The data, imitation, and Sim2Real chapter already identifies the risk that an old action chunk continues after the environment changes. This fixture supplies a smaller question to test before any implementation decision: for one specified perturbation and horizon, did replay amplify the nominal difference, and was offline replanning demonstrably available? It complements the broader policy-evaluation workshop, whose Sim2Real gate remains in force.

The durable output is a review packet, not a benchmark score: the task contract, paired traces, declared thresholds, unavailable branches, stop records, and the next question. If the packet cannot separate a measured author result from an unexecuted local exercise, or cannot show the branch conditions, it is not ready to guide an action-chunk change.

MENTAL MODEL / COORDINATES

The same point has different coordinates in different frames.

Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.

x = cos θ
y = sin θ

This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.

xy(0.87, 0.50)

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Measuring the Stability Assumption Behind Action Chunking ↗arxiv.orgPublished: 2026-10-01 · Accessed: 2026-10-04
Saved in this browser only.