What the new paper actually measures
The first version of Measuring the Stability Assumption Behind Action Chunking was submitted to arXiv on October 1, 2026. The authors ask a narrow but useful question: after a small injected error in an action, does the state difference shrink or grow? They compare two counterfactual trajectories. In the open-loop branch, the remaining actions from a recorded chunk are replayed without replanning. In the closed-loop branch, the policy receives the perturbed state and replans on later steps. Their rate is fitted over a finite observation window, with states classified as stable, unstable, or unresolved when the evidence is not decisive.
Those are the authors’ measurements in their stated simulation benchmarks and policy conditions. They do not establish a best chunk size for every robot, show that a short chunk is always safer, or report a result for this curriculum. The paper itself treats the fitted rate as a finite-horizon summary and separates error propagation from task performance.
This distinction matters because “the policy emitted eight actions” mixes two questions: how much a first action can move the task away from its nominal path, and whether later observations can recover from that move. A smooth demonstration can conceal either failure. A replayed remainder may preserve an obsolete plan; a replanning policy may still fail if the observation, task state, or recovery data are insufficient.
An unexecuted offline fixture
The following is original educational design inspired by the paper’s measurement framing. It has not been run here. It uses only a simulator or already-recorded, non-actuating rollouts. It does not connect a robot, command motors, train a model, or grant permission for a hardware test.
- 1Offline task contract and action envelope
- 2nominal trajectory and action stamp
- 3bounded synthetic perturbation
- 1Nominal branch
- 2replay remaining stored actions
- 3open-loop deviation
- 1Perturbed state
- 2replan only in an offline simulator
- 3closed-loop deviation
- 1Paired traces
- 2evidence and permission review
- 3next offline question or stop
First write the task contract: state representation, action units and frame, action envelope, success condition, distance measure, and a named reviewer. Pick a stored trajectory and an action stamp. Make a nominal branch with the unmodified action. In a separate synthetic branch, change only one selected action by a predeclared bounded perturbation. The bound must be expressed in the same units and frame as the action; an unexplained “small” displacement is not reproducible.
For the open-loop comparison, replay the remaining stored actions in the simulator without calling the policy. For a closed-loop comparison, ask the policy for later actions only if a resettable simulator can supply the perturbed observations. If the data contain no valid simulator state, no observation model, or no policy access, record closed_loop_unavailable. Do not substitute a guess that replanning would recover.
Before looking at outcomes, choose a distance such as object-pose error, end-effector error, or a task-state distance, and choose a fixed horizon. Save the nominal and perturbed time series, the horizon, perturbation vector, simulator and policy versions, random seed, and any reset. Label the result contracting, expanding, or unresolved only under thresholds stated in the task contract. The label describes this trace under these conditions; it is not a general stability certificate.
Permission and stop are part of the result
The fixture should log four decisions beside its error curve:
| Question | Required record | Safe outcome when evidence is absent |
|---|---|---|
| Is the perturbation inside the offline action envelope? | units, frame, bound, reviewer | stop the case |
| Can the branch be reset and reproduced? | state snapshot, seed, simulator version | mark unresolved |
| May a closed-loop branch call the policy? | explicit offline permission and input contract | do not replan |
| Does any proposed follow-up leave the task boundary? | stop reason and owner | stop; no hardware escalation |
An expanding open-loop trace can motivate a review of observation timing or chunk horizon. It cannot automatically shorten a controller’s chunk, authorize a live trial, or prove that the closed-loop policy is unsafe. A contracting trace is equally limited: repeated perturbations, sensor delay, contact changes, and a different task may reverse the conclusion. Keep rejected and stopped cases; excluding them makes recovery look better than the available evidence supports.
How this extends the existing chapter
The data, imitation, and Sim2Real chapter already identifies the risk that an old action chunk continues after the environment changes. This fixture supplies a smaller question to test before any implementation decision: for one specified perturbation and horizon, did replay amplify the nominal difference, and was offline replanning demonstrably available? It complements the broader policy-evaluation workshop, whose Sim2Real gate remains in force.
The durable output is a review packet, not a benchmark score: the task contract, paired traces, declared thresholds, unavailable branches, stop records, and the next question. If the packet cannot separate a measured author result from an unexecuted local exercise, or cannot show the branch conditions, it is not ready to guide an action-chunk change.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01