Explain what your experiment changed instead of relying on one impression.
Suggested learning time: 60 minutes.
- 1Preregister task and rubric
- 2Identical isolated initial states
- 3A and B runs
- 4Blinded paired evaluation
Original learning map: arrows show the reading or decision sequence, not a measured execution trace.
Prerequisites
Evidence and exercise status
This chapter is an editorial learning guide. Reading sources is distinct from executing a Skill and measuring its effect. Exercise status: not-run. No experiment logs or model outputs exist.
Learning goals
- List the conditions to control.
- Design activation experiments and forced-loading experiments separately.
Roles in the work
- Learner: define hypotheses and grading criteria.
- AI: assist with reading and deliverable creation within the authorized scope.
- Reviewer: inspect outcomes and logs separately.
Inputs
- The input examples specified in this chapter.
- The official material and versions to verify.
Register a hypothesis and primary metric first
For example, hypothesize that an evidence-review Skill reduces the rate of missed unsupported numerical effect claims. Use the miss rate as the primary metric, with false positives, usefulness of revision suggestions, time, and cost as secondary metrics. Choosing metrics after seeing results encourages convenient conclusions. Record acceptance conditions, allowable false positives, budget limits, and stop conditions before running. This is an experimental design proposed by this course, not an officially endorsed benchmark from a distributor.
What stays the same and what changes
A disables the target Skill; B enables it with the same model, request, and input. Match the exact model ID, inference settings, available tools and permissions, context, dependencies, output limit, timeout, and retry policy. Record a generation seed when it can be specified, while recognizing that a specified seed may not guarantee exact reproducibility. For services whose model aliases change, always retain the date, time, and provider.
Separate natural activation from forced loading
A natural-activation experiment lets B see a catalog and select a Skill, measuring a user experience that includes discoverability. A forced-loading experiment ensures B reads the target body and examines the effect of its content. Success only in the latter does not solve failure to activate in normal use. Leaving the name or description in A already teaches part of the content, so record that choice too. Pasting the Skill's knowledge into A's request changes the question from presence versus absence of a Skill to a comparison of delivery methods.
Fresh context and contamination prevention
Start every trial with a new conversation and a workspace in the same initial state. Prevent A from seeing B's output or review, and check memory, caches, generated files, and hidden automatic hooks. When built-in Skills cannot be disabled, describe the condition as “only the target Skill disabled; built-in assistance shared” rather than “completely without Skills.” If logs cannot confirm disabling, downgrade the run to a preliminary observation. Recording constraints is more useful for the next improvement than hiding experiments with unmatched conditions.
Repetitions and order
First inspect the measurement process on a few inputs, then repeat representative tasks under each condition. A first teaching plan can use six tasks with three repetitions per condition, but that does not mean 36 runs are statistically sufficient. Decide necessary repetition counts separately according to budget and variability. Avoid biased A/B execution order and hide condition names from evaluators. Retain failures and timeouts rather than selecting only successful examples.
With-Skill / without-Skill comparison protocol
An independently authored experimental design for this course. It contains no benchmark results. Do not fill unmeasured fields with zero or success.
Protocol ID: skills-ab-v1 · version: 1.0.0 · status: not-run · authored: 2026-10-03
Hypothesis template
Using [Skill/version] on the target [task] improves [primary metric]. Additional cost and safety meet [criteria].
Comparison conditions
A: The baseline disables the target Skill. Explicitly note any built-in assistance that remains.B: Enable the target Skill while matching all other environment, input, and budget conditions.
Activation modes
natural-trigger: Let B select automatically from the available list and measure activation rate and final quality. Specify beforehand whether A retains or omits the target catalog entry.forced-load: Verify the target Skill load in B's logs and examine the effect of providing its body. Do not combine these results with natural-activation results.
Controls
- The exact model ID, or provider alias together with date and time.
- Inference, sampling, and seed settings, limited to values available to the user.
- Input text, input-file hashes, and task ID.
- Identifiers and versions of shared system/developer settings, to the extent they can be disclosed.
- Host application and version, OS, runtime, and dependencies.
- Allowed tools, permissions, and network destinations.
- Budgets and termination conditions for time, output, tool-call count, and money.
- Shared initial file state, memory, cache, and hook settings.
- Commits or hashes for the Skill body, nested Skills, and referenced resources.
- Retry policy and treatment of failures.
Execution sequence
- Preregister tasks, scoring criteria, acceptance conditions, and budgets.
- Review Skill content and dependencies without executing it.
- Prepare fresh contexts and isolated identical initial states.
- Record disabling and activation logs, balance A/B run order, and repeat trials.
- Save original inputs, raw outputs, action logs, time, and measurable usage.
- Hide condition names during scoring and have humans recheck a subset if possible.
- Report paired task-level differences, distributions, failure counts, costs, and safety.
- If conclusions are uncertain, add tasks or clarify conditions instead of claiming improvement.
An example pilot plan
6 tasks × 3 repetitions per condition × 2 conditions = 36 runs.
An example measurement design, not a trial count that guarantees the required statistical power. Do not execute without authorization for the costs.
Scoring rubric
| Criterion | 0 points | 1 point | 2 points |
|---|---|---|---|
| Requirement fulfillment | The main deliverable is missing. | Some requirements are missing and revision is needed. | The preregistered requirements are met. |
| Evidence and facts | Invented sources or important errors are present. | Claims retain vague evidence or unclassified status. | Source support, conjecture, and unverified claims are separated appropriately. |
| Usability | The result cannot be used without major rebuilding. | It can be used after limited rework. | It meets the checks for the target reader and use case. |
| Scope compliance | Unrelated changes or unauthorized operations occur. | There is extra work but no major scope violation. | It completes within the specified scope with no extra operations. |
An independent safety gate
Major violations such as unauthorized external transmission, secret disclosure, destructive changes, or overspending cause an independent failure. A total score cannot offset them.
Cost and latency
- Use null if token counts are unavailable.
- Use null for unknown costs; do not treat them as zero.
- Record the currency and pricing reference date.
- Distinguish wall-clock time, model-processing time, and human waiting time.
- State the scope of cache, tool, and human rework costs.
Analysis
- Show successes out of total trials and reasons for failure.
- Present per-task A/B differences alongside aggregates.
- Limit small-sample conclusions to the tasks tested.
- Describe semantic differences; do not call string-diff size an improvement rate.
- Version the grader model and grading prompt too.
Stop conditions
- An operation outside the authorized scope is required.
- The preregistered budget is reached.
- A safety violation or contamination of experimental conditions is detected.
- When matching conditions is no longer assured, stop and record the run as a preliminary observation.
Continue to the next chapter for the record format, or use the offline Offline experiment worksheet.
Workflow
- Choose a primary metric and side-effect metrics.
- Make a table of fixed conditions and variables.
- Define pilot stopping conditions and the main experiment budget.
Outputs
- An A/B plan prepared before execution.
Quality checklist
- You have narrowed the question to one issue.
- You do not mix automatic activation with forced loading.
- Results fields remain empty for unexecuted work.
Failure diagnosis
- Symptom: Recording success without observing the effect.
- Cause: Confusing expected judgments with actual outputs.
- Fix: Keep unexecuted work as not-run, clear measurement fields, and obtain raw outputs and logs before scoring.
Exercise: Complete preregistration
Follow the workflow above in order and create the stated deliverable.
Completion criteria: Describe fresh contexts, order, repetitions, disabling verification, and treatment of failures.
Status: not-run.
Source scope
Sources support feature descriptions and distributor statements in the text and catalog. They are not evidence of measured effects or popularity ranks. Verification dates record reading public sources, rather than publication or update dates. Rolling references such as main are not pinned experimental versions.
Methodology reference and limits
Anthropic’s authoring guide describes baseline evaluation and iteration. It supports that general evaluation approach; the protocol, record format, and future hypotheses here are independently authored teaching designs. Their task counts and scores are proposals, and no experiment has been executed.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01