kumyu.Learn
← Kumyu home日本語 ↗

REGISTERED TAG

Model training and adaptation

Adjust learned model parameters and compare adaptation methods.

Published Updated
Intermediate

Image production flow: Reference, character consistency, Control, LoRA, Inpainting

Instead of relying on one-shot generation, we separate the responsibilities of reference, control, local correction, and LoRA when necessary, and transform it into a reproducible image production process.

15 min↗
Published Updated
Beginner

Map AI, machine learning and LLMs

Match the problem you want to solve to a mechanism, rather than treating AI as one uniform tool.

11 min↗
Published Updated
Beginner

Learn the basics of training, loss and evaluation

Treat inference and training as different computations, and design evaluation before training.

8 min↗
Published Updated
Intermediate

Choose a GPU platform and cost model

Align runtime, storage, stopping and data-transfer assumptions before comparing prices.

12 min↗
Published Updated
Intermediate

Change behavior with SFT, LoRA and QLoRA

Organize additional training by its purpose and its implementation method.

11 min↗
Published Updated
Beginner

Choose continued pretraining and preference optimization

Identify the goal and explain why a more demanding training method is needed.

12 min↗
Published Updated
Beginner

Adapt to Japanese: vocabulary, notation and tasks

Evaluate Japanese language ability separately from domain expertise.

9 min↗
Published Updated
Intermediate

Handle medical terminology: sources, negation, units and human review

Distinguish terminology lookup and document processing prototypes from validated diagnostic use.

12 min↗
Published Updated
Intermediate

Read DeepSeek-R1: what reasoning RL changes and what it does not promise

Read the 2025 DeepSeek-R1 paper through reward design, distillation, evaluation conditions, reproduction, and operational limits.

10 min↗
Published Updated
Intermediate

Read s1: what to measure when improving reasoning with little data

Translate the 2025 s1 paper into implementation decisions through test-time scaling, budget forcing, data selection, reproduction, and limits.

9 min↗
Published Updated
Intermediate

Read DAPO: do not view LLM reinforcement learning as an algorithm alone

Read the 2025 DAPO paper through reward, length bias, asynchronous system design, reproduction, and reward hacking.

9 min↗
Published Updated
Intermediate

Read 2025 reasoning-RL and efficiency papers into implementation

Read four 2025 papers on test-time compute allocation, RL training efficiency, shorter reasoning, and dynamic length control by separating claims from constraints.

10 min↗
Published Updated
Intermediate

2025 VLA map: read SmolVLA, π0/π0.5, OpenVLA, and GR00T using the same yardstick

Updated notes that interpret each presentation in terms of openness, body, data, behavioral expression, and evaluation boundaries, rather than ranking each presentation in a comparison table.

12 min↗
Published Updated
Intermediate

LeRobot practice: Make dataset, imitation learning, evaluation, and Sim2Real into one quality loop

Starting with LeRobot's data format and unified interface, learn the points of contact that are likely to fail from collection to actual machine evaluation.

15 min↗
Published Updated
Intermediate

Reading VLA papers to implementable designs: Comparative reading exercise of π0, SmolVLA, and GR00T

The novelty of a paper is not confused with performance ranking, but is broken down into behavioral expression, data, body, evaluation, and unverified boundaries.

16 min↗
Published Updated
Intermediate

Physical AI workshop: from data splits to policy evaluation and Sim2Real decisions

Turn public or synthetic episodes into reproducible splits, baseline and adaptation evaluations, failure analysis, and a justified decision not to proceed to hardware.

12 min↗