Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Classify failure
  2. 2prompt, retrieval or training hypothesis
  3. 3data and budget
  4. 4fixed evaluation
  5. 5adopt or revisit
Consider the sequence and each role.

Before you begin

Prerequisite: Change behavior with SFT, LoRA and QLoRA. Estimated study time: 50 minutes.

This chapter is an editorial guide. The exercise is UNRUN (not-run); it asks for a method selection plan rather than an executed training experiment.

What you will learn

  • Distinguish RAG, SFT, continued pretraining and preference optimization by their input data.
  • Create an experiment plan that includes regression and the cost of preparing data.

Continued pretraining adapts to a distribution of documents

Continued pretraining takes an existing model and continues training tasks such as next-token prediction using additional text. It can be considered when adapting to the vocabulary, writing style and distribution of a large collection of domain documents. Its purpose and data differ from SFT, which teaches instruction following through questions and example answers.

Providing correct domain documents does not guarantee that the model can accurately cite their facts whenever asked. Outdated knowledge, memorization of training data and regression in general capabilities can still occur. First check whether RAG or small-scale SFT can solve the problem. Then establish the required quantity of data, usage rights, computing budget and evaluations outside the target domain.

Preference optimization learns which answer is preferred

Preference optimization methods such as DPO use pairs of chosen and rejected answers to the same prompt. They can adjust policies such as politeness, concision and how to respond when evidence is insufficient. If the definition of a good answer is unclear, that ambiguity becomes part of the training.

Watch for evaluator biases, such as always favoring long answers or particular phrases. A preferred answer is not necessarily factually correct. Check for overlap between SFT data, preference data and evaluation data. For a beginner's first exercise, stop at explaining and designing the method; do not rush into complex reward or preference training.

See Transformers causal language modeling for next-token training, TRL SFTTrainer for supervised examples and TRL DPOTrainer’s dataset section for chosen/rejected completions. These describe distinct objectives and data contracts; none makes preference a factuality guarantee.

Test the hypothesis with the smallest change

If the problem is missing the latest information, consider updating documents and using RAG. If the model does not follow JSON requirements, consider schema validation, prompts and SFT. If the vocabulary distribution is substantially different, analyze the tokenizer and consider adaptation to a domain corpus. If answer policies are inconsistent, clarify the scoring criteria and consider preference optimization.

Every method needs a baseline model, the same evaluation, regression checks and reproducible settings. An experiment can still be valuable when it does not improve results, provided it helps locate the cause in retrieval, data quality, model capacity, inference settings or the evaluation itself.

Compare methods by purpose, data and cost

Method Purpose Input data Changes weights? Evaluate Cost considerations
Prompts / ordinary code Improve instructions, formats and validation. Instructions and a schema. No. Format compliance and task completion. A lower-cost starting point.
RAG Use external material, updates and sources. Documents the user may access and an index. No. Retrieval recall, citation agreement and abstention. Index updates, retrieval and generation.
SFT Teach examples of desired behavior. Inputs and example outputs. Yes. Task accuracy, format and regression. Data creation and training.
LoRA / QLoRA An implementation method that reduces the memory burden of updates. A base model and adapter training data. Yes. The evaluation appropriate to the training purpose above. A candidate for lighter updates than full-weight training; activations and other memory costs remain.
Continued pretraining Adapt to the distribution of domain text. A corpus with verified usage rights. Yes. Domain tasks and regression in general capabilities. Data preparation and computing can become substantial.
Preference optimization, such as DPO Adjust the desired answer policy. prompt, chosen and rejected answers. Yes. Independent preference evaluation and factuality. Comparative evaluation data and training.

The table includes both training purposes and weight update methods. Keep those two axes distinct when choosing a combination.

Roles and input → process → output

Role Responsibility
Learner Design the hypothesis, data and evaluation.
Model Attempt the specified transformation.
Application Enforce limits, validation and permissions.
Stage What it contains
Input A failure classification, available data, evaluation criteria and computing budget.
Process Start with small changes, choose a method and run a controlled comparison.
Output The reason for adopting a method, expected improvement, regression criteria and stopping conditions.

Workflow

  1. Narrow the investigation to one kind of failure.
  2. Use a non-training remedy as the baseline.
  3. Estimate the types and quantity of data needed.
  4. Prepare evaluations both within and outside the target task.
  5. If there is no improvement, return to the earlier hypothesis.

Exercise: assign a method to four problems

Status: UNRUN (not-run). Choose methods for four tasks: obtaining the latest specification, following a JSON format, adopting a domain writing style and controlling answer length.

Deliverable: a table mapping each method to its input data, improvement metric and risks.

Completion check: use separate columns for the kind of training and the weight update method.

Quality checklist

  • Can you distinguish RAG, SFT, continued pretraining and preference optimization by their inputs?
  • Does your experiment plan include regression and data costs?
  • Are the training purpose and weight update method shown separately?

Pitfalls and failure diagnosis

Do not use DPO as a fact checker. Continued pretraining alone does not establish that a system can provide safe medical support.

Caveat What to check
Treating DPO as a verifier of facts. Evaluate factual support separately from preference and writing style.
Assuming continued pretraining alone makes medical support safe. Use the terminology task’s source, negation and unit checks; keep medical judgments subject to human review.

Use the experiment worksheet to plan or record this experiment.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Transformers causal language modeling ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
TRL DPOTrainer ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
03
TRL SFTTrainer ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
04
PEFT quantization ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.