R1 mattered for more than displaying long thought
The DeepSeek-R1 paper, published 2025-01-22, brought wide attention to extracting reasoning behavior from language models with reinforcement learning. Its design question is not simply that a model writes a long explanation and succeeds at math or code. When a reward is placed on a machine-checkable result, what search behavior does the model learn, and what side effects follow? Making answers long by itself increases both cost and opportunities for misunderstanding.
- 1Verifiable problem
- 2multiple rollouts
- 3correctness and format reward
- 4policy update
- 1Changed reasoning behavior
- 2evaluation set
- 3distillation or production budget control
R1-Zero applies large-scale RL without a cold-start supervised fine-tuning phase. Tasks with programmatically checkable outcomes—final math answers and code tests—make rewards tractable. Multiple solutions for a question can yield relative signals for updates, allowing search, self-checking, and revision-like behavior to emerge without labeling every good explanation. The paper also reports that this path can produce unreadable prose, language mixing, and unstable format. R1 therefore combines limited cold-start data, reasoning-oriented RL, and stages for refusal and format. It is not a claim that RL alone makes a universal intelligence: verifiable outcomes, initial distribution, reward-hacking controls, and readable output all matter.
Separate the mechanism into three parts
First is outcome reward: a final numeric value or unit test can score a result without judging how persuasive its prose seems. Second is within-group relative assessment: taking several samples for one problem and updating with their differences can avoid separately training a value function. Third are constraints for format, language, and safety. Rewarding outcome alone leaves paths such as parser tricks, unnecessary verbosity, or output that cannot be explained to a reader.
- 1Problem
- 2N answers
- 3verifier
- 4reward group
- 1Reward group
- 2relative advantage
- 3update
- 1Correct result with format violation
- 2format reward or validator
- 3fail
Do not confuse showing a reasoning trace with proving reasoning is correct. What users need may be source documents, checked calculations, tool results, and reproducible steps, not a plausible long narrative. Keep the answer verifier separate from the verifier for user-facing explanation.
Distillation is an implementation option, not a free transfer
The paper reports distilling large-R1 reasoning outputs into smaller models. It does not guarantee that an expensive teacher’s behavior copies cleanly into a small environment. Traces contain teacher-specific templates and tokens irrelevant to the correct result. Evaluate a student through verified final answers and data including short counterexamples, rather than treating trace memorization as success. The official repository is an entry point for related models and materials; inspect licenses and revisions for code, weights, paper, and derived checkpoints separately.
Reproduction plan: do not start with large-scale RL
- Build about 50 verifiable cases: final arithmetic values, sandboxed code tests, or structured extraction checked against schema and references. Do not transmit production data.
- From a base model, collect multiple samples per question and record pass@1, pass@k, average tokens, and format-failure rate with fixed temperature, seed, and maximum tokens.
- Before reproducing RL, compare reuse of correct samples, rejection sampling, and short SFT. Small improvements are valuable when their cause is identifiable.
- If running RL, measure rollouts, verifier, advantage, update, and checkpoints separately. Stop if reward is always zero or always maximal.
- Seek failures with paraphrase, misleading instructions, long input, and format constraints; exclude gains that merely exploit the reward function.
Verifiable reward is powerful, and verifier holes are powerful too. Incomplete code tests invite patches that pass tests while breaking untested behavior. A final-answer-only math reward can confuse lucky answers and invalid format. Multiple samples can raise accuracy while increasing latency and cost. Long reasoning can repeat confidential input or make wrong intermediate work look trustworthy. The transferable lesson is to find a narrow machine-verifiable job and separate generation, scoring, policy updates, and user display. When an experiment fails, retain hypotheses about the verifier, saturated base model, insufficient samples, context pressure, or domain mismatch. Stop and inspect outputs when reward rises while human quality falls.
Separate a visible trace from a verifiable result
The 2025 Reddit reaction to R1 distills included reports of long output and inconsistent instruction following. Those anecdotes are not comparative measurements. They point to a useful evaluation split: score the final answer with a task verifier, then separately score format compliance, latency, token use, and harmful or unsupported claims. A correct answer delivered after an unusable delay is not interchangeable with a short verified answer.
For a local reproduction, retain the exact model revision and serving stack. Change only one factor—sampling, max tokens, or prompt template—per run, and compare paired tasks rather than averages from different prompt sets.
DeepSeekMath (2024) paired domain-targeted continued pretraining with GRPO before the later R1-style reasoning discussion. This makes the lineage a hypothesis about training components, not a proof that a later model’s behavior follows from one algorithm. Hold data mixture and evaluation set fixed before attributing a gain to RL alone.
SpecScale (submitted 2026-09-30) treats fine-grained verification as a way to prune a test-time search space. Its empirical results remain the paper authors’ results. The operational lesson is to distinguish “more branches” from “more verified progress”: log each candidate, verifier decision, deduplication event, and deferral; reject a budget increase when it only multiplies equivalent candidates or unverifiable claims.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
SOURCES
01YOUR NOTES