The useful question is which revision may persist
Robot improvement often gets described as a better model. Reconstruct, Practice, Go Real (RPG), submitted on October 1, asks a different question: how can an execution system improve its symbolic skill library and its system prompt while model weights stay fixed? Its contribution is not that an agent may edit code. It assigns different evidence to the roles that propose an edit, diagnose a failure, and decide whether a shared edit may remain.
This is a useful complement to the 2024 Lifelong Robot Library Learning preprint, which also grew a composable skill library through simulation. RPG narrows the operational problem further: a candidate repair is a change to shared infrastructure, so it must be checked beyond the task that suggested it. That is a lineage for the idea of a growing library, not evidence that the two systems have the same data, robot, or result.
- 1Offline trajectories + task descriptions
- 2construct a fixed simulation practice suite
- 1Runtime agent: robot-visible observations
- 2calls symbolic skills
- 3execution trace
- 1Privileged agent: same skills + simulator state
- 2reference trace for diagnosis only
- 1Traces + fixed evaluator
- 2candidate skill or prompt revision
- 3cross-task gate
- 1Accepted simulation revision
- 2frozen deployment candidate
- 3separate hardware safety review
The diagram is an original reading model. “Privileged” is especially important: it is a diagnosis aid in simulation, not information a deployed robot can assume it will receive. A library can contain a useful routine such as “place a tall object,” but its routine is only a program and a contract. The task evaluator, starting-state distribution, sensor data, control interface, and stop policy determine whether that routine means anything in a particular environment.
Separate the roles before calling anything an improvement
RPG builds practice tasks from an offline dataset and freezes their tasks and outcome checkers before practice. Its Runtime Agent selects and composes reusable skills using observations available to the robot. A separate Privileged Agent uses the same model and skills, but also sees simulator state. A Video Analyzer compares their traces with dataset video when available; an Implementor proposes one isolated revision; a Merger combines eligible revisions. Those names describe the authors’ experimental roles, not a production access-control design.
The important boundary is the data path:
| Role | May use during simulation practice | Must not become a deployment assumption |
|---|---|---|
| Runtime agent | robot-visible observations, skill return values, execution history | hidden object pose or a task evaluator’s internal state |
| Privileged agent | the same inputs plus simulator state, to make a diagnostic comparison | simulator state on a physical robot |
| Analyzer / implementor | recorded traces and a frozen task contract | authority to move hardware or silently accept its own edit |
| Cross-task gate | complete task outcomes from the fixed suite | a safety certification or a proof of real-world generalization |
This division answers a practical debugging question. If the privileged trace succeeds while the runtime trace fails, the missing information or perception interface is a candidate cause. If both fail, a shared skill or task contract may be implicated. Neither comparison proves causality by itself. It creates a narrower inspection target than asking a general-purpose model to “make the robot better.”
The paper’s reported gate is deliberately concrete: it evaluates a candidate across the practice suite, requires an increase in mean success, and rejects a candidate if any task loses more than one success in five fixed development trials. It then evaluates the merged revision again. That is an author-reported regression rule under its own simulator, seeds, evaluators, models, and budgets. It is not a universal acceptance threshold, and it does not replace a safety case.
What the reported results do and do not establish
The paper reports that its final simulated system achieved 209 successes in 220 held-out episodes over 22 tasks after fifteen practice rounds. It also reports 30 successful physical trials across three tasks after a shared calibration and hardware-adaptation procedure. These are author measurements, not an independent replication or a result for this curriculum’s robots.
Several details limit a transfer claim. Practice relied on privileged simulator state for one agent, while the deployment runtime did not. The physical evaluation used a YAM robot, fixed task criteria, ten trials per task, a stated reasoning/call budget, manual recreation of initial arrangements, and the authors’ calibration and adaptation procedure. The project page labels code as “coming soon”; no runnable repository, software license, weights, dataset license, controller setup, or safety approval was established in this review. A video or a perfect-looking tally cannot fill those gaps.
The paper also separates a system-level curve from a library-only comparison. That distinction is worth retaining in any future work: prompt edits, perception usage, model calls, skills, and task distributions can all move at once. A change in overall task success does not identify which component caused it. The authors’ own held-out curve is not strictly monotonic, which is a useful reminder that a development gate can miss later regressions.
An unexecuted N=1 worksheet: one revision, one offline boundary
Do not start by trying to reproduce RPG or by commanding a robot. Use a single already-recorded, non-actuating simulation episode or a synthetic trace. This worksheet is original educational design and was not run for this article.
- Write one task contract: initial-state identifier, allowed observation fields, symbolic skill name, success predicate, and a stop owner. Mark every field that comes from hidden simulator state; the Runtime column may not read it.
- Choose one proposed change, such as adding
verify_graspafter a pick. Preserve a baseline trace and a candidate trace under the same saved initial state. Do not change the prompt, evaluator, model version, or action interface at the same time. - Record four rows: event index, runtime-visible observation, skill call and return status, and evaluator outcome. A separate diagnostic note may quote hidden simulator state, but it must never enter the candidate’s runtime input.
- Set the only pass question in advance: did the candidate preserve the baseline predicate and improve the named failure without a new violation in this one trace? The only valid outputs are
candidate_supported_for_offline_review,rejected, orunknown. - Stop if the trace lacks an evaluator, the state cannot be reset, the change needs a privileged field at runtime, or a result would be used to justify hardware. Preserve that stopped record; do not retry with a different environment or relax the predicate.
This worksheet cannot establish the paper’s cross-task result, a safe grasp, simulator-to-robot transfer, or permission to actuate hardware. It is valuable because it makes the hidden-state boundary and the proposed revision visible before a team grants a program more reach.
A deployment decision needs another owner
RPG freezes its execution system before its physical evaluation; it does not use its privileged agent or video analyzer while the robot is moving. Keep that separation in a real program, but add controls the paper’s evaluation cannot grant: an approved operating envelope, emergency stop ownership, force and speed limits, exclusion zones, reset protocol, observation-loss behavior, and an accountable operator. The evaluator that accepts a simulated task is not the person or system allowed to expose a robot to people or property.
For a shared skill library, the durable artifact is therefore not “a robot that learned.” It is a revision packet: the base version, one proposed delta, task and evaluator versions, runtime-versus-privileged inputs, traces, regression results, unresolved failures, and the explicit decision that hardware remains out of scope. That packet lets a later reviewer decide whether the next work is better data, a better simulator contract, a code review, or a separately approved safety test.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01