Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Desired behavior examples
  2. 2SFT objective
  3. 3full-weight or LoRA updates
  4. 4held-out task evaluation
  1. 1Quantized base model
  2. 2QLoRA adapter training
  3. 3base plus adapter at inference
Consider the sequence and each role.

Before you begin

Prerequisite: RAG: retrieve knowledge instead of packing it into weights. Estimated study time: 50 minutes.

This chapter is an editorial guide. All exercises and code examples are UNRUN (not-run). Their presence does not establish that training completed or that quality improved.

What you will learn

  • Explain why SFT and LoRA are not competing alternatives.
  • Define the inputs, outputs and stopping conditions of a small adapter experiment.

Separate the purpose from the method

Supervised fine-tuning, or SFT, adjusts behavior using examples of desired inputs and outputs. LoRA limits the updates to small, low-rank matrices; SFT can be performed with LoRA. SFT can also update all the weights. QLoRA trains an adapter while using a quantized base model. It saves memory, but this does not mean every tensor becomes 4-bit.

RAG provides access to source material. SFT is better understood as a way to adjust output formats, terminology use and task procedures through examples. Using SFT only to memorize an FAQ makes updates and source tracking difficult. RAG and SFT can be combined when the task requires both.

The LoRA paper describes low-rank updates; the QLoRA paper describes adapter training through a frozen quantized base. PEFT’s quantization guide documents a supported library path. MLX LM’s separate LoRA/QLoRA guide is the reference for the Apple Silicon commands below.

Even a small dataset needs a design

Begin with short, fictional examples you write yourself to check that the training path works. Include examples that abstain when a question has no answer, boundary cases for the format, and inputs with spelling mistakes. Adding only positive examples may reinforce the habit of making definite claims in the requested format even when information is missing.

Separate train, validation and test data at the level of the original problem. Rephrasing the same question and putting its variants into different splits creates leakage. A small toy experiment is not proof of a quality improvement. Run the same fixed evaluation on the baseline model and the model with the adapter, looking both for the intended improvement and for regression in general capabilities.

Keep Mac and CUDA training paths distinct

On a Mac, MLX LM's support for LoRA and quantized models is one candidate. In a CUDA environment, Transformers, PEFT and TRL are another combination to consider. Do not assume an example using bitsandbytes runs unchanged on a Mac. Check the hardware and quantization backend support of the specific versions you use.

The command below is a smoke test using a small MLX-converted model, batch size 1, four trainable layers and ten iterations. Those settings are not a recommended training budget for proving effectiveness. Before running it, check that every example is within 512 tokens after applying the chat template. Do not use an external experiment tracker for this exercise. Even ten iterations do not come with a guarantee of processing time or memory use.

Roles and input → process → output

Role Responsibility
Learner Design the hypothesis, data and evaluation.
Model Attempt the specified transformation.
Application Enforce limits, validation and permissions.
Stage What it contains
Input A base model, train/validation/test JSONL files and training settings.
Process Keep the base model and update the adapter; compare using a fixed evaluation.
Output Adapter configuration and weights, logs, the base model reference and comparison results.

Workflow

  1. Save the evaluation results before training.
  2. Check your own dataset for duplicates and usage rights.
  3. Check the token limit using short examples.
  4. Try only a few iterations on a small model.
  5. Increase the data and training budget only after confirming an improvement.

UNRUN example: check the training path with prepared toy data

Status: UNRUN (not-run). Prepare toy-data/train.jsonl, valid.jsonl and test.jsonl manually first. Every example must be at most 512 tokens after applying the template and must contain no personal information. Ten iterations are a smoke test. Installing dependencies and obtaining the model for the first time are separate actions that this command would trigger. These published commands have not been run.

The Japanese prompt is retained exactly because it is part of the experiment input.

python -m pip install "mlx-lm[train]"
mlx_lm.lora --help
mlx_lm.lora --model mlx-community/Qwen2.5-0.5B-Instruct-4bit --train --data ./toy-data --iters 10 --batch-size 1 --num-layers 4 --adapter-path ./toy-adapter
mlx_lm.generate --model mlx-community/Qwen2.5-0.5B-Instruct-4bit --adapter-path ./toy-adapter --prompt "架空の商品について回答してください。" --max-tokens 128

UNRUN example: one JSONL record

Status: UNRUN (not-run). Use one example per line. Do not use data from real patients or customers. The original Japanese target strings stay unchanged.

{"messages": [{"role": "user", "content": "資料: 青いノートは80ページ。質問: 何ページですか。JSONで回答してください。"}, {"role": "assistant", "content": "{\"answer\":\"80ページ\",\"evidence\":\"青いノートは80ページ\",\"unknown\":false}"}]}

Exercise: structure fictional product inquiries

Status: UNRUN (not-run). Create a small dataset that converts inquiries about fictional products into JSON with answer, evidence and unknown fields.

Deliverable: the split JSONL files, plus before-and-after measurements of format compliance, agreement with evidence and abstention rate.

Completion check: improved formatting alone must not be described as increased knowledge.

Quality checklist

  • Can you explain why SFT and LoRA are not competing alternatives?
  • Have you defined the inputs, outputs and stopping conditions of the small adapter experiment?
  • Are you avoiding a claim that the model gained knowledge merely because its format improved?

Pitfalls and failure diagnosis

An adapter alone cannot perform inference. A falling loss is not the same as an improvement for the intended use. Do not apply an adapter to a different base model revision without validation.

Caveat What to check
Expecting an adapter to perform inference on its own. Retain the required base model, revision, target modules and tokenizer.
Treating lower loss as proof of improvement for the task. Compare the adapted model with the baseline on the same held-out tasks and inspect regressions.
Applying an adapter to a different base model revision without checking it. Verify the base revision and rerun the fixed evaluation before using the adapter.

Use the experiment worksheet to plan or record this experiment.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
LoRA original paper ↗arxiv.orgPublished: Unknown · Accessed: 2026-10-03
02
QLoRA original paper ↗arxiv.orgPublished: Unknown · Accessed: 2026-10-03
03
PEFT quantization ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
04
TRL SFTTrainer ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
05
MLX LM LoRA/QLoRA training guidePublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.