Version teaching material, Skills and experiments separately.
Suggested learning time: 60 minutes.
- 1Changed model or Skill
- 2Impact review
- 3Same regression cases
- 4Keep, revise, or remove
Original learning map: arrows show the reading or decision sequence, not a measured execution trace.
Prerequisites
Evidence and exercise status
This chapter is an editorial learning guide. Reading sources is distinct from executing a Skill and measuring its effect. Exercise status: not-run. No experiment logs or model outputs exist.
Learning goals
- Manage versions, references, and regression tests.
- Decide when to retire or reduce a Skill.
Roles in the work
- Learner: define hypotheses and grading criteria.
- AI: assist with reading and deliverable creation within the authorized scope.
- Reviewer: inspect outcomes and logs separately.
Inputs
- The input examples specified in this chapter.
- The official material and versions to verify.
Keep three kinds of versions separately
At minimum, version the teaching material, Skill, and experiment separately. Correcting a textbook typo does not change an experimental condition. A single-word change in a Skill body, updated references, a model switch, or an execution-environment change can alter behavior. To continue a comparison, do not update all of them at once. Append new run IDs to the ledger instead of overwriting old records, and retain the versions you can restore.
Priorities for reviewing changes
On an update, inspect changes to external destinations, permissions, scripts, dependencies, hooks, and licenses before looking at new capabilities. Next read activation descriptions, output formats, and failure behavior. A new Skill can override conventions in an older project, so include representative existing workflows in regression tests. Compare the benefits of pinning versions against the risk of missed updates instead of choosing automatic updates solely for convenience.
Keep maintenance small
Rather than arranging three similar Skills side by side, separate their roles and make names and descriptions clear. Combining company-specific rules, general procedures, and links to current specifications in one document mixes reasons for updating it. Separate materials with different update frequencies or owners. Repeatedly generated helper code may be shared as a reviewed script, but adds dependency and safety maintenance. Writing procedures down does not eliminate maintenance.
Removing a Skill can be a success
If removing a Skill preserves quality while reducing time or failures with the same model and representative tasks, consider removal, shortening, or restricted activation. Being useful initially does not make it necessary forever. Conversely, a Skill that prevents rare costly failures should not be judged only by an average, even if everyday effects are hard to see. Base maintenance decisions on usage frequency, revision cost, false activation outside scope, and observed serious failures. Do not use the number of Skills as a success metric.
Workflow
- Choose five safety items to inspect on every change.
- Define positive and negative activation examples and quality-regression tasks.
- Write conditions for adoption, deferral, retaining an old version, and retirement.
Outputs
- A review sheet for Skill updates.
Quality checklist
- References are versioned too.
- Behavior changes are distinguished from textbook corrections.
- You have defined retirement conditions.
Failure diagnosis
- Symptom: Recording success without observing the effect.
- Cause: Confusing expected judgments with actual outputs.
- Fix: Keep unexecuted work as not-run, clear measurement fields, and obtain raw outputs and logs before scoring.
Exercise: Create an update gate
Follow the workflow above in order and create the stated deliverable.
Completion criteria: Include a policy against overwriting versions and stop conditions for permission changes.
Status: not-run.
Source scope
Sources support feature descriptions and distributor statements in the text and catalog. They are not evidence of measured effects or popularity ranks. Verification dates record reading public sources, rather than publication or update dates. Rolling references such as main are not pinned experimental versions.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01