Allow approximately 50 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.

Prerequisite: Read Hugging Face pages and model files

Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Choose CPU or MLX
  2. 2inspect dependencies and first download
  3. 3capped fictional input
  4. 4manual inference
  5. 5record output and unknowns
Consider the sequence and each role.

Learning objectives

  • Distinguish the CPU and MLX execution paths and select one to try.
  • Separate preparation from execution, with no automatic execution.

Responsibilities

  • Learner: design hypotheses, data and evaluation.
  • Model: try the specified transformation.
  • Application: enforce limits, validation and permissions.

Input → process → output

Input Process Output
Short fictional text and a checked small model Tokenize → load the model → infer with a generation limit Generated text, token counts, execution conditions and an experiment record that retains unknown measurements

Choose one path for the first run

On Apple Silicon, one path uses MLX LM. To inspect the concepts independently of OS, choose the Transformers CPU path. Ollama is a candidate for easier model management and local APIs; llama.cpp is a candidate for detailed control over GGUF and inference settings. You do not need to install all four at the beginning.

The following examples are unrun teaching code. Displaying this page installs nothing. Check that Python and the runtime are supported, verify the distribution source and license, and review free disk space and network access before manually performing only the steps you select. The first from_pretrained/load operation downloads the model.

The CPU example follows the Qwen model card’s chat-template and generation path, with an explicit CPU choice and shorter generation cap. The community conversion card identifies the MLX model. MLX LM’s official usage reference supplies the CLI; this course has not downloaded or run either model.

Check with short input and output

Begin with short text without personal information, such as a fictional product description. The CPU example is restricted to a 0.5B model, at most 512 input tokens and 64 output tokens. It places the entire model on CPU and avoids unexpected device assignment through device_map=auto. Its purpose is to inspect the execution path rather than compete on speed.

The MLX example uses a 4-bit conversion of the same small model and limits output to 128 tokens. Observe memory and execution time even when the input is short. Stop if it does not finish, memory pressure worsens or the application becomes unresponsive; do not move to a larger model.

Define success narrowly

Success at this stage means recording the model ID and settings, receiving a short generation and observing resource consumption. Whether the answer is correct, adequate for Japanese tasks or able to serve several users belongs to the next evaluation. Do not label unrun code as verified on hardware.

If you expose a local API, initially bind it only to 127.0.0.1. Internet or LAN exposure, unauthenticated access and sending data to an external logging service are outside the first exercise. Even for a local application, separately review cloud features and telemetry settings.

Code examples to try manually

These examples are unrun. Reading or copying does not execute them. Some commands install dependencies or download models. Choose a path after reviewing environment compatibility, capacity and licenses. The Japanese benchmark inputs below intentionally remain unchanged in both editions.

Environment preparation: CPU path only

Status: not run (not-run).

python3 -m venv .venv
source .venv/bin/activate
python -m pip install torch transformers
python -m pip freeze > environment-cpu.txt

This installs dependencies from official PyPI. Record versions when you run it, then pin a verified environment for reproducibility.

Minimal Transformers inference on CPU

Status: not run (not-run).

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "Qwen/Qwen2.5-0.5B-Instruct"
torch.set_num_threads(4)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=False)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=False).to("cpu")
model.eval()
messages = [{"role": "user", "content": "架空の製品説明です。青いノートは80ページです。一文で要約してください。"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt")
assert inputs["input_ids"].shape[-1] <= 512, "入力が長すぎます"
with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=64, do_sample=False)
new_tokens = output[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
print({"input_tokens": inputs["input_ids"].shape[-1], "output_tokens": len(new_tokens)})

CPU floating-point behavior and optimizations depend on library versions. Do not treat the current main revision as a reproducible version; pin a revision after the initial verification.

Apple Silicon: MLX path only

Status: not run (not-run).

python3 -m venv .venv-mlx
source .venv-mlx/bin/activate
python -m pip install mlx-lm
mlx_lm.generate --help
mlx_lm.generate --model mlx-community/Qwen2.5-0.5B-Instruct-4bit --prompt "青いノートは80ページです。一文で要約してください。" --max-tokens 128
python -m pip freeze > environment-mlx.txt

A supported Apple Silicon environment is required. The public model ID causes an initial download. The page does not automatically execute these commands.

Follow the workflow

  1. Check the supported official Python environment and OS.
  2. Create a dedicated virtual environment.
  3. Choose either CPU or MLX and install that path’s dependencies.
  4. Read the code, understand the initial download, then execute it manually.
  5. Save the package listing and model revision.

Review quality

  • Distinguish the CPU and MLX execution paths and select one to try.
  • Separate preparation from execution, with no automatic execution.
  • Have you recorded “it ran” and “it was accurate” in separate fields?

Diagnose failures

Caveat What to check
These examples were not executed while creating this textbook. Keep the examples marked not-run until you execute them and retain the environment and output.
Do not respond to errors by adding unknown installation scripts or remote model code. Check the fixed model and official runtime requirements; inspect unfamiliar code before allowing it to execute.

Exercise: Run a small LLM locally

Status: not run (not-run).

Summarize the same short text, then manually check whether the model added facts or omitted too much.

Deliverable: The original text, output, model information, settings and observation notes.

Completion check: Have you recorded “it ran” and “it was accurate” in separate fields?

Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Qwen2.5-0.5B-Instruct model card ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
MLX community conversion: Qwen2.5-0.5B-Instruct-4bit ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
03
MLX LM official repositoryPublished: Unknown · Accessed: 2026-10-03
04
Transformers Auto classes ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
05
Ollama FAQ ↗docs.ollama.comPublished: Unknown · Accessed: 2026-10-03
06
llama.cpp official repositoryPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.