The unit under test is a configured system

At 2026-10-01 12:55:07 UTC (21:55:07 JST), Agents Are Systems, Not Models was submitted to arXiv as v1. It treats an agent as a model plus the harness that supplies context, tools, time, and feedback. That framing changes a practical question. “Did model X solve this?” is too small when the task information, prompt, tool contract, stop rule, and verification authority also changed.

In four scientific workflows, the authors report that about 54% of outcome variance among genuine attempts came from rerunning the same configuration rather than changing a configuration axis. This is an author result for their task suite and harness, not a general variance rate for every agent, model, or coding task. Its useful implication is modest: repeat a frozen setup before crediting a visible score difference to a model or prompt.

The July 2024 AI Agents That Matter abstract had already argued for separating accuracy from cost, downstream requirements, and reproducibility. It does not establish this paper's repeat-variance result. It supplies a narrower historical question: if an agent result cannot be reproduced under its stated conditions, an accuracy value alone is not enough to guide system selection.

  1. 1Task contract
  2. 2frozen configuration manifest
  3. 3repeated offline attempts
  1. 1Each attempt
  2. 2completion, verifier result, time, cost, tool errors
  1. 1Same configuration
  2. 2spread and failure categories
  3. 3keep, revise one axis, or stop
  1. 1Verifier authority
  2. 2explicitly bounded evidence
  3. 3never permission to act outside the fixture
Consider the sequence and each role.

The paper compares information, reasoning, prompted self-verification, time budget, and backbone model. It also finds, under its conditions, that a dedicated verification tool changes verification behavior more than merely asking for verification. This does not mean that every tool is trustworthy. A verifier can be incomplete, stale, over-permissive, or authorized only for a toy task.

An unexecuted N=1 repeat worksheet

This is an original, unexecuted offline fixture. It does not call a provider API, install a runtime, run a benchmark, access private files, or perform an external action. Pick one synthetic task with a deterministic checker, such as producing a JSON object that satisfies a local schema and a two-case arithmetic consistency check.

Before any attempt, write one manifest:

Field Lock it to one concrete value
Task and inputs one synthetic task ID and immutable input file hash
System model identifier, harness revision, prompt/instruction revision, tool list
Configuration axes reasoning mode, maximum steps/time/tokens, context bundle, temperature/seed if available
Verifier checker revision, exact predicates, unavailable cases, and its authority scope
Accounting start/end timestamp, token or request count if exposed, wall time, local cost estimate method

Run no more than three offline attempts under the identical manifest. For each, retain the raw output, checker result, stop reason, tool-call/error count, wall time, and any available cost denominator. Do not average away a parser error, timeout, refusal, or checker exception. State unknown when the account/runtime does not expose a cost or token denominator; do not invent one from a price page or another model.

Only after the repeats are recorded may one axis change. For example, add a local checker that validates the JSON shape, while keeping the task, model revision, prompt, and budget unchanged. The checker can answer whether that fixture's predicates pass. It cannot authorize deployment, sending a message, changing cloud resources, using another account's data, or claiming that the task is true outside its synthetic input.

Read the spread before tuning a score

Use a small table, not a conclusion from the best run.

Outcome pattern Diagnosis to record Next safe step
All attempts fail the same predicate task information, capability, or checker may be wrong inspect the contract; do not add budget automatically
Results differ under the same manifest stochasticity or hidden state is plausible preserve all runs; investigate state/version leakage
Checker passes but a human rule fails checker has an authority gap reject the checker as a release gate and narrow the claim
Time or cost is missing accounting boundary is absent retain unknown; do not compare efficiency

The paper's public repository was readable on October 4, but its root contained only a README and .gitignore, and its API metadata reported no license at that observation. Treat the paper's reported benchmark/trajectory release as a source to inspect, not as proof that a runnable, licensed reproduction is currently available. The paper itself limits its conclusions to four tasks in two scientific domains and leaves memory and multi-agent coordination unvaried.

This fixture complements the existing paper-reading workflow: a paper score becomes a local hypothesis only after the task contract, system revision, repeat count, verifier boundary, and missing denominators are visible. A successful synthetic row is evidence about that row; it is never authority for a production action.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Agents Are Systems, Not Models: Rethinking Agentic Evaluation ↗arxiv.orgPublished: 2026-10-01 · Accessed: 2026-10-04
02
Rethinking agent evaluation repositoryPublished: Unknown · Accessed: 2026-10-04
03
AI Agents That Matter ↗arxiv.orgPublished: 2024-07-01 · Accessed: 2026-10-04
Saved in this browser only.