The question changes when the output is an artifact
An answer benchmark can grade a string. An embodied-design evaluation must grade something that was built: a morphology, its controller, and the strategy that makes them work together. That does not make the result a hardware result. It changes the unit of evidence from what the model said to what an executable artifact did under a declared simulator contract.
The October 5, 2026 arXiv v1 paper ArtifactArena evaluates generated robot morphology, controller, and strategy in a MuJoCo contest. Its author measurements support simulator-specific design and control claims; they do not establish real-world robot capability. No result was independently reproduced here.
- 1Task rules + fixed simulator
- 2model produces morphology + controller + strategy
- 1Validator
- 2executable artifact
- 3matched opponents and declared seeds
- 4outcomes
- 1Outcome log
- 2separate design quality from simulator contract and test budget
- 1Any real robot claim
- 2separate safety review, authorization, and physical evidence
This is an original teaching diagram. The paper supplies the arena premise; the comparison protocol and claim boundaries below are a reusable evaluation plan.
A useful lineage: construction is not yet competition
In 2024, MMLU-Pro evaluated selected multiple-choice answers. BuildArena then evaluated generated physical constructions; ArtifactArena evaluates generated artifacts in competition. The scores cannot be compared across those tasks: each names a different output and evaluator.
In October 2025, BuildArena framed language-guided construction under physics-based checks. ArtifactArena changes the task: generated artifacts compete in a common arena. A bridge that supports a load and a robot that defeats an opponent require different contracts; neither proves manufacture, reliability, or field safety.
The simulator is also a contract, not a neutral photograph of the world. MuJoCo is a general-purpose physics engine for articulated structures and contact. Its inputs, numerical settings, contact model, and observation interface determine what is measured. Changing any of these while comparing models can turn a design comparison into a simulator comparison.
The nearby VLA patch-defense lesson asks whether an intervention preserves clean policy behavior. This chapter asks a different question: given an artifact generator, did the generated body, controller, and strategy earn its result under one fixed evaluator?
What must be held equal
The ArtifactArena harness material exposes why “we tested it in a simulator” is too weak. A submission has MJCF robot structure and a controller; the arena owns global physics settings and validates submissions. That interface is one part of the comparison contract.
Before comparing generator A with generator B, freeze this table. Do not choose a winner from an aggregate score if any row differs without an explicit reason.
| Comparison field | Record | Why it can reverse a conclusion |
|---|---|---|
| Arena contract | simulator version, rules, timestep, contact settings, validator revision | a changed evaluator changes the task |
| Artifact interface | allowed morphology primitives, controller API, observation schema | one generator may receive an easier design space |
| Budget | model/checkpoint, prompt, API-call or token budget, wall-clock cap | more search is not necessarily better design |
| Match set | opponent snapshot, pairings, order, side/start state, seed list | a design can exploit one incumbent opponent |
| Selection | qualification rule, candidate-selection procedure, retries | selecting after many failures leaks search effort |
| Outcome | win/loss/draw, validation error, crash, timeout, NaN/Inf, disqualification | a win rate without failures hides fragility |
Use the same seed list for every candidate, but do not stop there. Keep a held-out seed list and a held-out opponent snapshot. The first detects a design that only works at familiar starts; the second detects a strategy tuned to a known opponent. If a new arena revision changes the leaderboard population, mark it as a new evaluation rather than appending its wins to an old total.
An offline fixture for a fair comparison
No model, simulator, GPU, API, or robot is run in this exercise. Make a small review fixture with two fictional artifacts, A and B, and two rows per artifact:
run_id,artifact_id,arena_rev,opponent_rev,seed,budget_calls,result,failure_code
r01,A,arena-01,opponents-01,11,40,win,none
r02,A,arena-01,opponents-01,29,40,loss,none
r03,B,arena-01,opponents-01,11,40,invalid,xml_validation
r04,B,arena-01,opponents-01,29,40,timeout,timeoutAdd the morphology hash, controller hash, prompt version, model/checkpoint identifier, candidate-selection rule, and log hash. Then answer four questions before calculating any rate.
loss means a competitive loss only. Keep timeout as its own result and invalid as its own validation failure.
- Did both artifacts receive the same arena, opponents, seeds, and budget?
- Are
invalidandtimeoutcounted separately from a competitivelossrather than discarded? - Can a reviewer reproduce exactly which artifact was selected from a search run?
- Would the conclusion survive if the held-out seeds and opponents replaced the familiar set?
An apparent improvement that vanishes on the held-out slice is a failure diagnosis, not an embarrassment. Label it opponent_overfit, seed_sensitivity, budget_asymmetry, validator_failure, or unknown; preserve the trace. Do not repair the record by rerunning only favorable seeds.
Three generation modes are not three deployment permissions
Treat the three modes described in the public harness repository as a comparison variable: sampling makes candidates without iterative feedback; verifier-grounded refinement uses verdicts; Design Lab expands tool-assisted search. More feedback may improve a submitted artifact but also changes access, retries, and selection. Name every result model + checkpoint + harness + budget + arena revision; otherwise a model claim silently includes its search wrapper.
Cost, permission, and the boundary at reality
The sources do not provide a reader-specific price or a guaranteed reproducible compute bill. An actual run may require a model provider or local checkpoint, simulator dependencies, storage for artifacts and traces, and enough compute for repeated matches. Record the actual price, hardware, and time only after executing a defined workload; do not estimate a production budget from a paper leaderboard.
The paper links the ArtifactArena organization. Its public harness-public and design-lab-harness repositories declare MIT licenses. Their current READMEs specify Python 3.10 and MuJoCo 3.10; model runs can require provider keys, a Hugging Face dataset token, and—in Design Lab—Linux isolation features. These are requirements, not a cost estimate. They do not grant a reader API, dataset, model-weight, or hardware permission. The arXiv version identifies CC BY-NC-SA 4.0 for paper text; rights for generated artifacts remain unverified.
Finally, do not turn a simulated robot into a physical one because it won. A physical trial needs separately authorized hardware, a reviewed safety plan, emergency-stop ownership, speed/force limits, a containment area, supervision, reset and recovery rules, and evidence from the actual hardware. None was requested or performed here.
The mental model is simple: an artifact score is evidence about a particular generator inside a declared arena. It becomes a useful result only when the arena, budget, seeds, opponents, and failures remain inspectable.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01