REGISTERED TAG
Model training and adaptation
Adjust learned model parameters and compare adaptation methods.
Image production flow: Reference, character consistency, Control, LoRA, Inpainting
Instead of relying on one-shot generation, we separate the responsibilities of reference, control, local correction, and LoRA when necessary, and transform it into a reproducible image production process.
Map AI, machine learning and LLMs
Match the problem you want to solve to a mechanism, rather than treating AI as one uniform tool.
Learn the basics of training, loss and evaluation
Treat inference and training as different computations, and design evaluation before training.
Choose a GPU platform and cost model
Align runtime, storage, stopping and data-transfer assumptions before comparing prices.
Change behavior with SFT, LoRA and QLoRA
Organize additional training by its purpose and its implementation method.
Choose continued pretraining and preference optimization
Identify the goal and explain why a more demanding training method is needed.
Adapt to Japanese: vocabulary, notation and tasks
Evaluate Japanese language ability separately from domain expertise.
Handle medical terminology: sources, negation, units and human review
Distinguish terminology lookup and document processing prototypes from validated diagnostic use.
Read DeepSeek-R1: what reasoning RL changes and what it does not promise
Read the 2025 DeepSeek-R1 paper through reward design, distillation, evaluation conditions, reproduction, and operational limits.
Read s1: what to measure when improving reasoning with little data
Translate the 2025 s1 paper into implementation decisions through test-time scaling, budget forcing, data selection, reproduction, and limits.
Read DAPO: do not view LLM reinforcement learning as an algorithm alone
Read the 2025 DAPO paper through reward, length bias, asynchronous system design, reproduction, and reward hacking.
Read 2025 reasoning-RL and efficiency papers into implementation
Read four 2025 papers on test-time compute allocation, RL training efficiency, shorter reasoning, and dynamic length control by separating claims from constraints.
2025 VLA map: read SmolVLA, π0/π0.5, OpenVLA, and GR00T using the same yardstick
Updated notes that interpret each presentation in terms of openness, body, data, behavioral expression, and evaluation boundaries, rather than ranking each presentation in a comparison table.
LeRobot practice: Make dataset, imitation learning, evaluation, and Sim2Real into one quality loop
Starting with LeRobot's data format and unified interface, learn the points of contact that are likely to fail from collection to actual machine evaluation.
Reading VLA papers to implementable designs: Comparative reading exercise of π0, SmolVLA, and GR00T
The novelty of a paper is not confused with performance ranking, but is broken down into behavioral expression, data, body, evaluation, and unverified boundaries.
Physical AI workshop: from data splits to policy evaluation and Sim2Real decisions
Turn public or synthetic episodes into reproducible splits, baseline and adaptation evaluations, failure analysis, and a justified decision not to proceed to hardware.