Production fails on the remaining one case, not the average score
A demo that succeeds nine times in ten can look compelling. In production, the remaining time can mean a wrong charge, secret disclosure, missing draft, or stopped work. An average benchmark does not say on which screen that happens or to whom. Move evaluation from “model rank” to “did the user’s work finish safely?”
- 1User goal
- 2input classification
- 3evidence retrieval
- 4model generation
- 1Generation
- 2schema and authority check
- 3display or confirmation
- 4execute
- 1Every stage
- 2trace, cost, and failure
- 3return to evaluation set
Separate ambiguous decisions a model may assist with—wording, search queries, classification—from deterministic decisions the application must own: payment, deletion, permission changes, email recipient, and publication. Giving the second group to generated text erases the accountability boundary regardless of prompt quality.
Build an initial evaluation set of 20–50 cases from past requests, manual procedures, and prohibited errors, not necessarily 1,000 questions. Fix input, allowed evidence, expected shape, prohibitions, and scoring method. Do not send personal or secret production data to outside evaluation without authorization; anonymization can remain re-identifiable, so include synthetic cases. Use provider guidance such as OpenAI safety practices, Anthropic tool use, Google Gemini API documentation, and AI Gateway as primary starting points, then test the actual workflow.
For every request, attach an identity, request ID, model/revision, retrieval IDs, tool intent, result, stop reason, latency, and bounded usage record. Log raw content only under a purpose and retention policy. Validate tool arguments against schema and current authorization, use allowlists and confirmation for side effects, make writes idempotent or deduplicated, and offer recovery when a stream or connection ends. Set ceilings by user, feature, and day; retries need a count, backoff, and total budget. Ship one version behind feature flags, compare it against a control on fixed cases, and keep a rollback target. A safety system is not a refusal string: it is evidence, authority, action control, monitoring, and recovery that can be tested end to end.
Incident-ready evaluation
Write expected handling before the incident: no evidence should produce a hold rather than invented support; a tool timeout should leave an action visibly pending or safely failed; a duplicated request should yield one recorded effect; a lost connection should not erase a draft. Add these cases to the same evaluation suite as normal success. Review a sample of traces after release for unsupported claims, wrong-user access, unexpected tool calls, budget anomalies, and changes in refusal behavior. Classify the finding by retrieval, model, validator, authorization, UI, or provider operation so the repair lands in the responsible layer.
Re-run the locked suite for prompt, model, retrieval, validator, tool, and policy changes. A feature flag is useful only if it has a named owner, measured release condition, and a tested disable path. Retain the prior configuration and a small before/after result record. This gives an operator a concrete way to stop harm without guessing which of many moving AI parts changed.
2024–2026 change: evaluation includes adversarial and economic behavior
A 2024 prompt-quality score is inadequate for a 2026 system that can browse, call tools, or process customer context. Recent primary reports from OpenAI (2026-09-30) and Anthropic (2026-09-10) describe their own security work. They are useful threat inputs, not a substitute for testing an application's controls.
Extend the fixed evaluation set with control cases: tool arguments that name another tenant, an instruction hidden in retrieved data, replay of a request ID, excessive output, upstream quota exhaustion, and a user who revokes consent during a job. Score the system on safe terminal state as well as answer quality: denied without side effect, held for review, completed once, or explicitly failed. Make cost a request-level invariant—maximum input bytes, tools, retries, elapsed time, and spend—not an after-the-fact dashboard. Preserve enough redacted trace data to reproduce a failure without retaining secrets or full personal content.
A September 29 r/ClaudeAI incident discussion is a community-maintained incident thread, not a service-level source. It motivates an outage fixture—stable user-visible failure, retry ceiling, and no duplicate side effect—but cannot establish an SLA or control effectiveness.
Anthropic’s October 2024 computer-use beta explicitly described the capability as experimental and error-prone. It made action side effects a first-class control problem rather than an output-quality problem. Keep a 2026 regression fixture where a visible action is interrupted after intent selection but before execution; the required result is one recorded safe terminal state and no repeated side effect.
MENTAL MODEL / VERIFICATION COST
The value of a decision depends on downstream work.
Assumptions: one second for the judgment, half of the candidates retained, and equal verification time. Full parallelism needs enough compute and concurrency. Compare success rate and total cost, including wrong judgments and retries. These figures are estimates, not measurements.
SOURCES
01YOUR NOTES