A paper’s largest number is not a design answer
An abstract says briefly what improved. An implementer needs the conditions, comparison, cost, and trade-off that produced that improvement. “Higher accuracy” means something different if it came from more inference compute, different training data, or a leaky evaluation. The purpose of reading is not memorizing conclusions; it is constructing a hypothesis that can transfer to your problem and a boundary for what cannot.
- 1Abstract claim
- 2problem and baseline
- 3method change
- 4conditions and evaluation
- 1Matching conditions
- 2minimal reproduction
- 3refute or adopt on your data
- 1Different conditions
- 2record as a hypothesis
- 3do not make an implementation conclusion
Compress title and abstract into one sentence: who changed what to address which failure, measured how? Then seek evidence that could break the sentence. Was the baseline strong, was compute equal, were there multiple seeds, was the test set repeatedly tuned, and are failures shown? Looking for disproof first prevents an attractive figure from owning the design.
Keep four layers in separate notes
The problem layer names operational pain: costly long reasoning, unstable generation, or retrieval without evidence. The method layer states whether loss, data, architecture, sampling, or system implementation changed. The evaluation layer holds datasets, metrics, controls, compute, inference budget, and statistics. The operations layer holds required GPUs, latency, memory, licensing, monitoring, and failure behavior.
Papers can detail the method and leave operations thin because the research scope differs. Measure missing operational layers yourself. arXiv versions can change, so retain the cited version and date according to its versioning guidance. A repository with the same name as a paper is not automatically the experiment commit.
A baseline is a measuring instrument
When a proposed method wins, ask whether its baseline was tuned comparably. Learning rate, batch, output length, prompt, retriever, and budget can give one side an advantage. For LLMs, equal model size is insufficient: input tokens, output tokens, sample count, tool calls, wall time, and GPU type all affect inference compute.
- 1Proposed score
- 2check equal budget
- 3record a mismatch as a condition difference
- 1Equal budget
- 2inspect failures and variance
- 3small reproduction
- 1Reproduction gap
- 2implementation, data, or randomness difference
- 3update hypothesis
You need not retrain an entire paper. Begin with an author checkpoint or small setting and reproduce one table row or one failure example. Preserve environment, commit, seed, hardware, data revision, and run duration. The reproducibility checklist helps identify what was not recorded. A non-reproduction is useful evidence about what transfer requires.
Do not convert a number into a product promise
Benchmark accuracy does not simultaneously mean user satisfaction, legal correctness, low toxicity, or an SLA. Data contamination, selection bias, evaluator differences, out-of-distribution inputs, and language differences matter. Aggregators can help discovery, but use a source such as the Papers with Code data repository to return to the original paper; never retain only a screenshot of a number.
Translate work into acceptance criteria: “summarize Japanese internal policy with citations; an unsupported assertion fails,” or “return a draft within 15 seconds and preserve it for retry on timeout.” If the paper’s task differs from that criterion, retain it as reference rather than adoption evidence.
Practice: a two-hour paper card
- Record bibliography, arXiv version, author repository, license, and access date.
- Write one sentence each for problem, change, baseline, main figure, figure conditions, and author-stated limitations.
- List two matching and two differing conditions in your work. Hold implementation when differences are material.
- Choose one minimal reproduction condition: a public example validates a schema, identical input returns a cited passage, or compute can be measured.
- Label the result reproduced, partially reproduced, unexecuted, or refuted. Reading alone is not verification.
Make the failure table central. Preserve stated limitations, weaker datasets, added compute, and settings you could not reproduce. If the paper does not state a limitation, write “not found in the paper”; do not invent one. Keep reading notes (citations and interpretation) separate from experiment notes (commands, environment, results, failure). When a result differs, retain dataset revision, random state, evaluation code, and why a run stopped rather than keeping only the best attempt.
Treat version dates as an experimental variable
For preprints, record the arXiv version number and date separately from the first submission date. For code, record commit SHA, dependency lock, hardware, and evaluation harness revision. A claim may remain the same while the implementation, benchmark protocol, or appendix changes. The reproduction target is the tuple, not the title.
Build a negative-control run before adoption: preserve the baseline and deliberately remove the claimed mechanism while holding compute and data fixed. If the reported gain remains, inspect leakage, tuning, or the harness before attributing it to the method.
Quiet-STaR (2024) is a useful reading exercise because its abstract identifies both a mechanism and compute constraints. Convert neither into a product claim: extract the stated evaluation, make a small fixture that measures task outcome and added compute, and write the failure condition before choosing an implementation.
A September 2026 paper on budgeted multi-attribute verification makes the evaluation cost itself part of the method. Read it as a checklist prompt, not a certified policy: identify the attributes a verifier checks, the cost of each check, and the consequence of abstaining. For a reproduction note, predeclare a verification budget and record which claims were checked, deferred, or left unsupported.
MENTAL MODEL / REASONING ORDER
From an announcement to your own decision.
Compare the announcement with the conditions in the paper and official documentation.
SOURCES
01YOUR NOTES