Allow approximately 35 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.
Prerequisite: Understand models, tokens and inference
Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.
- 1Split by original document
- 2training data
- 3weight updates
- 4independent evaluation
- 1Held-out test data
- 2fixed scoring
- 3improvement and regression
Learning objectives
- Explain the roles of forward computation, loss, backward computation and the optimizer.
- Separate train, validation and test data without leakage.
Responsibilities
- Learner: design hypotheses, data and evaluation.
- Model: try the specified transformation.
- Application: enforce limits, validation and permissions.
Input → process → output
| Input | Process | Output |
|---|---|---|
| Training examples you may use and independent evaluation examples | Prediction → loss → gradients → update → fixed evaluation | A checkpoint or adapter, training logs and task-specific evaluation results |
Training quantifies errors and adjusts weights
Supervised learning uses pairs of inputs and correct answers. A forward pass makes predictions, a loss quantifies their difference from the correct answers, a backward pass computes gradients for the weights, and an optimizer updates those weights. LLM pretraining predominantly uses self-supervised learning in which the next token in text becomes the target answer.
Inference usually does not retain gradients or optimizer state. Training needs both of those, together with intermediate activations. “The model file fits into memory” therefore does not mean “this model can be trained in that memory.”
The TRL SFTTrainer documentation describes the trainer’s datasets and loss configuration. The MLX LM LoRA/QLoRA training guide is the separate Apple Silicon implementation reference; its author-reported examples are not measurements for this course.
Turn basic terms into experimental controls
An epoch describes how many passes are made through the complete dataset. Batch size is the number of examples processed at once. Learning rate controls the size of an update. Gradient accumulation collects gradients across multiple small batches before an update, increasing the effective batch without holding more examples simultaneously. Recording a seed does not guarantee identical results across different devices and implementations.
A typical sign of overfitting is decreasing training loss while performance on unseen examples gets worse. Before choosing a more complex model, inspecting a few examples, removing duplicates and clarifying ambiguous labels can often produce more useful improvements.
Fix the evaluation set before starting
Use train data for learning, validation data for choosing settings, and test data for the final comparison. Randomly splitting document fragments can place the same document on both sides and overestimate performance. Split at a unit that matches what will be unknown in production, such as original documents, users or time periods.
Loss and perplexity measure prediction as a language model; they do not directly prove safe business use. Measure format compliance, agreement with evidence, missing requirements, abstention when an answer is wrong, and processing time separately. Use model-based scoring as a supplement, and consider grader bias and expert verification.
Follow the workflow
- Define failure categories appropriate to the use case.
- Split data at the original-document level.
- Save baseline performance before training.
- Change only one setting at a time.
- Inspect regressions as well as improvements.
Review quality
- Explain the roles of forward computation, loss, backward computation and the optimizer.
- Separate train, validation and test data without leakage.
- Have you avoided treating the results of 20 questions as a general performance guarantee?
Diagnose failures
| Caveat | What to check |
|---|---|
| Do not prove improvement using the same questions that appear in training data. | Split by original document and check for overlap before evaluating on held-out data. |
| Do not assert that a small difference on a small evaluation set is statistically reliable. | Report the evaluation-set size and failure cases alongside the difference. |
Exercise: Learn the basics of training, loss and evaluation
Status: not run (not-run).
Design a small 20-question test with four dimensions: output format, factual agreement, abstention and speed.
Deliverable: An evaluation set with correct answers, scoring criteria and exclusion conditions.
Completion check: Have you avoided treating the results of 20 questions as a general performance guarantee?
Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01