Record comparison conditions, failure cases and scoring reasons.

Suggested learning time: 60 minutes.

  1. 1Raw inputs and outputs
  2. 2Observed metrics or null
  3. 3Scoring reasons
  4. 4Redacted public record
Consider the sequence and each role.

Original learning map: arrows show the reading or decision sequence, not a measured execution trace.

Prerequisites

Evidence and exercise status

This chapter is an editorial learning guide. Reading sources is distinct from executing a Skill and measuring its effect. Exercise status: not-run. No experiment logs or model outputs exist.

Learning goals

  • Record quality, time, cost, and safety separately.
  • Read results without hiding small samples or missing measurements.

Roles in the work

  • Learner: define hypotheses and grading criteria.
  • AI: assist with reading and deliverable creation within the authorized scope.
  • Reviewer: inspect outcomes and logs separately.

Inputs

  • The input examples specified in this chapter.
  • The official material and versions to verify.

Do not rush to a single overall score

Score requirement fulfillment, accuracy of evidence, usability, unnecessary changes, and permission compliance separately. For each 0–2-point criterion, create concrete anchors: 0 means missing or wrong, 1 means partially satisfied but needing revision, and 2 means meeting the specified standard. Length and jargon counts are not quality itself. For code, read the diff in addition to running tests; for teaching material, check sources and reader comprehension. Do not offset serious safety violations with a high total score.

Report only measured cost and time

Save input tokens, output tokens, cache treatment, and tool charges separately when available. If a subscription service does not disclose per-trial cost, use null rather than assuming zero. Wall-clock time from starting to obtaining a deliverable differs from model-processing time and time waiting for human confirmation. Stating the scope lets you compare a slower process with less rework against a quick process needing later correction. Measure human revision time separately when possible.

Read output differences from four viewpoints

First examine added and missing content; second, fact corrections and new errors; third, structure and usability; fourth, behavioral differences. A string diff can be large just because paragraphs moved, so its size is not an improvement rate. Write observations such as “B has a field indicating missing sources; A does not,” then separately interpret that the Skill's procedure may have contributed. Do not assert a cause that the experimental design has not isolated.

Retain averages together with failure cases

Examine paired A/B differences for each task before reporting means, medians, variability, and successes out of total trials. With few trials, show n explicitly and do not generalize to broad use cases. If using statistical interval estimation, retain the method and assumptions. If another model grades outputs, also record the grader model, prompt, anonymization, order, and any regrading. A model's score is not an absolute human judgment.

Separate archival records from public records

An archival ledger links inputs, raw outputs, grading reasons, logs, and model and Skill versions. A public record removes personal information and secrets and includes only publishable deliverables. Preserve the original logs and create a separate edited public version. Distinguish not-run, running, failed, and completed. Even when plans and results share a screen, labels and empty fields must prevent readers from confusing them.

Complete template for an unexecuted record

This is a record-format example, not an experimental result. null means unknown or not obtained; do not replace it with zero or success. Its status remains not-run.

Field What to retain
task Task ID, original prompt, input files, and input hash.
model Provider, exact model ID, snapshot, alias status, settings, and unavailable settings.
runtime Client/version/OS, tools/permissions/budget, fresh context, initial state, baseline isolation, built-in Skills, and cache.
skill Enabled state, repo/path/commit/hash, declared version, observed activation and evidence, resources, nested Skills, and hooks.
output Raw text, deliverable files, transcript reference, and output hash.
metrics Input/output/cached tokens, tool calls, wall/model/human-wait/rework time, cost, currency, and price date.
evaluation Blinded condition label, grader, grader model, rubric version, scores, safety violations, and evidence.
comparison Paired run, observed differences, interpretation, and uncertainties.
failure Failure information; null for an unexecuted record.
publication Required redaction, sanitized output reference, and permission to publish.
{
  "schemaVersion": "1.0.0",
  "experimentId": "skills-ab-example-not-run",
  "runId": null,
  "status": "not-run",
  "plannedAt": "2026-10-03",
  "startedAt": null,
  "endedAt": null,
  "condition": null,
  "mode": null,
  "task": {
    "id": null,
    "prompt": null,
    "inputFiles": [],
    "inputHash": null
  },
  "model": {
    "provider": null,
    "modelId": null,
    "snapshotId": null,
    "isAlias": null,
    "settings": {},
    "settingsUnavailable": []
  },
  "runtime": {
    "client": null,
    "version": null,
    "os": null,
    "tools": [],
    "permissions": [],
    "budget": {},
    "freshContextVerified": false,
    "sharedInitialStateHash": null,
    "baselineIsolationVerified": false,
    "knownBuiltInSkills": [],
    "cacheNotes": null
  },
  "skill": {
    "enabled": null,
    "repositoryUrl": null,
    "path": null,
    "commit": null,
    "contentHash": null,
    "declaredVersion": null,
    "activationObserved": null,
    "activationEvidence": null,
    "loadedResources": [],
    "nestedSkills": [],
    "hooks": []
  },
  "output": {
    "rawText": null,
    "files": [],
    "transcriptRef": null,
    "outputHash": null
  },
  "metrics": {
    "inputTokens": null,
    "outputTokens": null,
    "cachedTokens": null,
    "toolCalls": null,
    "wallClockMs": null,
    "modelMs": null,
    "humanWaitMs": null,
    "humanReworkMs": null,
    "cost": null,
    "currency": null,
    "priceDate": null,
    "measurementNotes": null
  },
  "evaluation": {
    "blindedLabel": null,
    "graderType": null,
    "graderModel": null,
    "rubricVersion": null,
    "scores": [],
    "safetyViolations": [],
    "evidence": []
  },
  "comparison": {
    "pairedRunId": null,
    "observedDifferences": [],
    "interpretation": null,
    "uncertainties": []
  },
  "failure": null,
  "publication": {
    "redactionRequired": true,
    "sanitizedOutputRef": null,
    "allowedToPublish": false
  }
}

Executed experiment count

The source experiments list is empty: []. There are zero executed comparison experiments and no recorded model outputs, times, token counts, costs, or measured effects. Preserve the empty fields above. Planning an exercise does not make it a measured run.

Offline experiment worksheet

Workflow

  1. Prepare model, Skill, and input identification fields in the supplied template.
  2. Make task-specific scoring criteria concrete.
  3. Save an unexecuted record and reread it to check that it cannot be mistaken for a result.

Outputs

  • An experiment-ledger template and scoring sheet.

Quality checklist

  • Inputs are paired with raw outputs.
  • Observations and interpretations are separate.
  • You do not hide safety violations inside an overall score.

Failure diagnosis

  • Symptom: Recording success without observing the effect.
  • Cause: Confusing expected judgments with actual outputs.
  • Fix: Keep unexecuted work as not-run, clear measurement fields, and obtain raw outputs and logs before scoring.

Exercise: Prepare an empty ledger without inventing values

Follow the workflow above in order and create the stated deliverable.

Completion criteria: Unknown measurements are null, unobtained outputs are empty, and no fabricated results are present.

Status: not-run.

Source scope

Sources support feature descriptions and distributor statements in the text and catalog. They are not evidence of measured effects or popularity ranks. Verification dates record reading public sources, rather than publication or update dates. Rolling references such as main are not pinned experimental versions.

Offline experiment worksheet

Methodology reference and limits

Anthropic’s authoring guide describes baseline evaluation and iteration. It supports that general evaluation approach; the protocol, record format, and future hypotheses here are independently authored teaching designs. Their task counts and scores are proposals, and no experiment has been executed.

Use the local experiment ledger

Open the browser-local experiment ledger. Start with a planned record. Unknown values remain blank, and an executed record requires saved output evidence. The ledger does not run models or upload records. You can keep using the worksheet above as a separate document.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Official documentationAnthropic: Skill authoring best practices — Evaluation and iteration ↗platform.claude.comPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.