Allow approximately 35 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.

Prerequisite: Map AI, machine learning and LLMs

Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Text plus chat template
  2. 2tokenizer
  3. 3input prefill
  4. 4repeated decode
  5. 5output token count
Consider the sequence and each role.

Learning objectives

  • Explain the roles of weights, the tokenizer, context and output length.
  • Record input and output token counts separately.

Responsibilities

  • Learner: design hypotheses, data and evaluation.
  • Model: try the specified transformation.
  • Application: enforce limits, validation and permissions.

Input → process → output

Input Process Output
A token sequence organized with the chat template Prefill → token selection → repeated decoding Generated tokens and the text decoded from them

Weights alone do not make a model run

A neural network model consists of a computational structure and learned numerical values called weights. To use it, you also need a tokenizer, configuration files, a chat template that arranges a conversation, and a runtime. Even when weights can be converted to another format, mismatches among these parts can prevent the expected response.

Tokens are not the same as characters or words. How Japanese, English, numbers, whitespace and symbols are divided depends on the tokenizer. Do not assume one Japanese character equals one token: count with the tokenizer of the model you use. An embedding represents tokens or other items as numerical vectors; it is not a human-readable dictionary of meaning.

Inference has two different timing phases

In a decoder LLM, prefill processes the input token sequence, then decode generates successive tokens. Time to first token, or TTFT, differs from the subsequent generation rate in tokens per second. Providing a long document can make prefill expensive even when the answer is short.

The context window contains input, conversation history, necessary special tokens and generated tokens. A long context listed on a model card does not guarantee comfortable use at that length on your machine. Manage inputs through summarization, retrieval and separation of history rather than adding conversation turns without a limit.

The Qwen example model card shows the chat-template path; the Transformers cache documentation explains why past keys and values are reused during decoding. Neither source measures this exercise on a reader’s machine.

The limits of probabilistic continuation

A trained LLM computes a distribution over the next token. Temperature and similar settings change token selection; they do not guarantee facts. Even greedy generation can produce incorrect content, and a plausible explanation is not proof of its basis. Design the system so calculators handle calculations, retrieval handles current information, and validated tools handle actions that need authorization.

A Base model provides a foundation for predicting continuations, while an Instruct model has been adapted for responding to instructions. Models of the same size can therefore require different usage. For the first conversational experiment, pair an Instruct model with its official chat template to reduce errors caused by template mismatch.

Follow the workflow

  1. Count the same text with the selected tokenizer.
  2. Set input and generation limits separately.
  3. Compare runs with and without conversation history.
  4. Record time to first token and generation speed separately.

Review quality

  • Explain the roles of weights, the tokenizer, context and output length.
  • Record input and output token counts separately.
  • Have you avoided generalizing this character-to-token ratio to other models?

Diagnose failures

Caveat What to check
Lowering temperature does not necessarily make answers accurate. Hold the task fixed and check factual support as well as output variation.
API compatibility does not imply equivalent model capability or compatible chat templates. Record the model, tokenizer and chat template before comparing the same requests.

Exercise: Understand models, tokens and inference

Status: not run (not-run).

Prepare a short Japanese sentence, an English sentence with the same meaning, and a JSON example. Record both token count and character count for each.

Deliverable: A token-count comparison with the exact model ID.

Completion check: Have you avoided generalizing this character-to-token ratio to other models?

Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Qwen2.5-0.5B-Instruct model card ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
Transformers KV cache ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
03
Transformers Auto classes ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.