LEARNING PATH
RESEARCH
Turn new mechanisms into experiments you can run.
Read AI papers into implementation decisions: claims, conditions, reproduction, and failure
A textbook for turning papers into practical experiments by separating a claimed improvement from its problem, conditions, reproduction, trade-offs, and failures.
Read DeepSeek-R1: what reasoning RL changes and what it does not promise
Read the 2025 DeepSeek-R1 paper through reward design, distillation, evaluation conditions, reproduction, and operational limits.
Read s1: what to measure when improving reasoning with little data
Translate the 2025 s1 paper into implementation decisions through test-time scaling, budget forcing, data selection, reproduction, and limits.
Read DAPO: do not view LLM reinforcement learning as an algorithm alone
Read the 2025 DAPO paper through reward, length bias, asynchronous system design, reproduction, and reward hacking.
Read 2025 reasoning-RL and efficiency papers into implementation
Read four 2025 papers on test-time compute allocation, RL training efficiency, shorter reasoning, and dynamic length control by separating claims from constraints.