Allow approximately 35 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.

Prerequisite: Estimate memory for 24GB and 32GB Macs

Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Original model card and license
  2. 2conversion publisher
  3. 3format and runtime compatibility
  4. 4pinned model record
Consider the sequence and each role.

Learning objectives

  • Record the model ID and revision.
  • Distinguish the roles of MLX formats, GGUF and safetensors.

Responsibilities

  • Learner: design hypotheses, data and evaluation.
  • Model: try the specified transformation.
  • Application: enforce limits, validation and permissions.

Input → process → output

Input Process Output
Model ID, model card, license and file listing Check provenance, compatibility, usage conditions and capacity A candidate-model record with a pinned revision

The Hub distributes models and records their context

Hugging Face Hub contains models, datasets and applications. Read more than a model page’s name: check publisher, original model, intended use, restrictions, license, required libraries and file sizes. Recording and pinning a verified commit revision makes comparisons more reproducible than continually using an evolving main branch.

The small examples in this textbook use Qwen/Qwen2.5-0.5B-Instruct and its community conversion, mlx-community/Qwen2.5-0.5B-Instruct-4bit. They are examples whose existence and published official usage have been checked, not claims that these are the latest or best models. The conversion’s publisher is different from the developer of the original model.

Compare the original Qwen card with the MLX community conversion. The conversion card identifies the source model and conversion with mlx-lm 0.18.1; that conversion tool version is historical provenance, not a required current runtime version. The Hub model-card guidance explains the role of license and intended-use metadata.

This course retains the Qwen2.5 series from September 2024 as a small, reproducible teaching candidate. The date comes from the model card’s series citation; it is historical context, not a claim that this is the newest model in October 2026. Choosing a later model still requires its own license, files, revision and evaluation. Qwen2.5 model card.

Match storage format to the runtime

safetensors is a tensor-storage format; its presence does not mean any runtime can run the model. Transformers reads the model architecture and configuration, MLX LM uses supported MLX formats, and llama.cpp uses compatible GGUF files. Renaming an extension does not convert a model.

A LoRA adapter is usually a small set of additional weights and cannot hold a conversation by itself. Keep the base model, revision, target modules and tokenizer it was built for. If you change the combination of quantized or unquantized weights and adapters, reevaluate it instead of reusing previous evaluation results.

License and download safety

Being public, being free to download and being usable in a commercial service are different properties. Read the conditions for the original model, derived model and dataset separately. Follow any redistribution requirements, attribution duties or usage restrictions; pause publication if the terms are unclear.

Do not casually set trust_remote_code=True for an unknown model. Custom code can execute during model loading. Preferring a safer storage format still does not guarantee that packages and runtime code are safe. Begin with official examples, known models and an isolated environment with ordinary privileges, and check the download size and destination.

Follow the workflow

  1. Choose an official or verifiable publisher.
  2. Trace the conversion back to its original model.
  3. Retain the license.
  4. Check compatibility between files and runtime.
  5. Record the revision and package versions in the experiment ledger.

Review quality

  • Record the model ID and revision.
  • Distinguish the roles of MLX formats, GGUF and safetensors.
  • Have you separated capability claims on the model card from your own measured results?

Diagnose failures

Caveat What to check
Do not treat open weights and unrestricted open source as synonyms. Read the original and converted model licenses, including redistribution and use restrictions.
Do not republish a third-party benchmark as the speed measured on your machine. Keep the provider’s benchmark separate; record your own hardware, runtime, inputs and results.

Exercise: Read Hugging Face pages and model files

Status: not run (not-run).

Compare the original example model with the MLX conversion. Record the different publishers, formats and whether quantization is used.

Deliverable: One model-selection card.

Completion check: Have you separated capability claims on the model card from your own measured results?

Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Hugging Face Model Cards ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
Qwen2.5-0.5B-Instruct model card ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
03
MLX community conversion: Qwen2.5-0.5B-Instruct-4bit ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
04
Transformers Auto classes ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
05
llama.cpp official repositoryPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.