The first three weights to separate

“A 7B model runs on a Mac” is only half true. Whether the weights fit, whether they continue to fit in a long conversation, and whether they return within the waiting time you need are distinct questions. Weights are mostly determined by parameter count and precision. Conversation working space grows in the KV cache. Latency has different behavior for reading the prompt and generating one token at a time. Collapsing all three into “required memory” makes it impossible to explain why a quantized model starts but fails on a long document, or why GPU utilization is high while the first response remains slow.

  1. 1Model weights
  2. 2resident at startup
  3. 3primarily reduced by quantization
  1. 1Input tokens
  2. 2KV cache
  3. 3grows with conversations and concurrency
  1. 1Prefill
  2. 2first token
  3. 3decode
  4. 4per-token waiting time
Consider the sequence and each role.

Weights are layer matrices. For FP16, parameter count times two bytes is a lower-bound estimate. Tensor layout, runtime buffers, and space for the OS and runtime are added in practice. Four-bit quantization is not a magic exact quarter: per-block scales and zero points, unquantized tensors, and activations remain. Still, generation is often bandwidth-bound, so smaller weights can improve both speed and fit.

Quantization is where you pay error, not simply a bit count

Quantization maps continuous weights into limited representations. A common approach finds a scale for a small group of weights, then approximates values with integers and that scale. If outliers share one scale, resolution for the remaining values becomes coarse; smaller groups increase metadata. Therefore a Q4 label alone cannot compare quality. Method, group size, calibration, target model, and evaluation task belong together.

Compare an FP16-equivalent and a candidate quantization with the same prompt, temperature, and output limit. “Does the prose sound natural?” is insufficient. Fix about 20 tasks you cannot afford to fail: numeric extraction, JSON validity, short code tests, and summaries containing proper nouns. Measure several runs and use a median after excluding cold start. Record model name and commit, quantization-file hash, context length, and input/output token counts; without them, next week’s result is not reproducible.

The KV cache is a table that avoids recomputation, not the conversation itself

At every token, a Transformer attends to prior tokens. Saving the Key and Value for each layer prevents recomputing all prior tokens for the next one. This stored table is the KV cache. Roughly, it grows with layer count, KV-head count, head dimension, token count, and bytes per element. Even when model weights are four-bit, a FP16/BF16 KV cache can dominate long documents or many sessions.

Specifying a long context does not always allocate all of that length immediately, but an implementation that permits the worst case often reserves heavily. A RAG design that always inserts whole large documents lowers concurrency before quality is even considered. Retrieve a small candidate set, pass quotation-sized units, and summarize conversation state where needed. This also makes evidence easier to trace.

The vLLM PagedAttention explanation describes allocating logical KV-cache blocks to physical blocks instead of one fixed contiguous area. Paging reduces unused gaps among requests of different lengths and makes shared prefixes easier to handle. It is a GPU-serving optimization, not an automatic choice for light solo interaction on a Mac.

MLX, llama.cpp, and vLLM are not competing for the identical job

For Apple Silicon, MLX is an array framework that can use unified memory and Metal, making research, conversion, and light inference approachable from Python. When using surrounding implementations such as MLX LM, verify model format, usage conditions, and quantization format individually. MLX itself does not grant a model-distribution license.

llama.cpp is an inference implementation centered on GGUF that can run across CPU, Metal, CUDA, and other environments. It suits trying one model from a terminal, running a local HTTP server, or keeping a distribution small. Speed varies substantially with backend, threads, GPU offload, batch size, and prompt length. Its README can establish available features; it does not guarantee a particular model’s quality or license.

vLLM is strong as a GPU server design for many requests. Continuous batching, KV-cache management, and an OpenAI-compatible server can justify it. It is not a product-name promise of maximum speed for one desktop user. CUDA-capable GPUs, drivers, model architecture, and operational monitoring are part of the decision. For a first learning machine, reduce the failure surface with llama.cpp or MLX; evaluate vLLM when multiple users or load testing actually require it.

  1. 1Personal experiment
  2. 2MLX or llama.cpp
  3. 3fix input, output, and quality evaluation
  1. 1Multiple requests and GPU operations
  2. 2vLLM
  3. 3observe concurrency, p95 latency, and KV use
  1. 1Quality regression
  2. 2move one quantization level back
  3. 3measure again on the same set
Consider the sequence and each role.

A minimal hands-on experiment

  1. Read the weights license and permitted uses at the model’s official distribution page. Do not conflate it with an OSS runtime license.
  2. Start with 20 short input/output cases in jsonl. Give every line an expected check: JSON parsing, a regular expression, a unit test, or human scoring.
  3. Run two quantizations of the same model with the same seed, temperature, output cap, and context length. Record startup, prefill, and decode separately.
  4. Observe actual memory use and failure around 4,000, 16,000, and 32,000-token-equivalent inputs. Recheck conclusions made with padded input against real documents.
  5. In the result table, write “adopted under these conditions” and “unverified beyond this conversation length.” Do not present one tokens-per-second figure as universal.

Many failures begin with an ambiguous measurement target rather than a weak model. Fast short chat and evidence-grounded reading of a long contract cannot be chosen by the same benchmark. Separating weights, KV cache, and input/output behavior reveals whether the next investment is RAM, VRAM, retrieval, or evaluation data.

September 2026 runtime check: releases can alter the measurement surface

The llama.cpp release stream showed September 2026 builds and platform assets. That is a current maintenance signal, not a throughput claim for your machine. Re-run the same prompt corpus after a runtime upgrade and record exact tag/commit, backend, context length, quantization, warm-up condition, prefill time, decode tokens per second, peak memory, and failures. A LocalLLaMA discussion about fragmented Apple Silicon optimization is community experience; it motivates measuring your stack rather than choosing from reported numbers.

The b10935 llama.cpp release, published on 2026-09-13, is a dated runtime artifact rather than a benchmark. Use its exact tag when repeating a local measurement; do not call a generic releases index a September release.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

SOURCES

01
llama.cpp README ↗github.com · unknown
02
MLX documentation ↗ml-explore.github.io · unknown
03
vLLM documentation: PagedAttention ↗docs.vllm.ai · unknown
04
llama.cpp release b10935 ↗github.com · 2026-09-13
05
LocalLLaMA discussion (community experience) ↗www.reddit.com · 2026-08-15

YOUR NOTES