Allow approximately 35 minutes. Exercise status: not run (not-run). This chapter is an editorial learning guide, not a record of measured operation or quality.

Prerequisite: Learn the basics of training, loss and evaluation

Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Physical capacity
  2. 2reserve OS and applications
  3. 3budget weights and KV cache
  4. 4small trial
  5. 5observe peak and stop conditions
Consider the sequence and each role.

Learning objectives

  • Estimate the size of 4-bit weights.
  • Choose a testing order and stopping conditions for generic 24GB and 32GB scenarios.

Responsibilities

  • Learner: design hypotheses, data and evaluation.
  • Model: try the specified transformation.
  • Application: enforce limits, validation and permissions.

Input → process → output

Input Process Output
Physical memory, model size, precision, context and concurrency Estimate each component, then measure the peak in a small experiment A memory budget and feasibility decision for each inference configuration

Do not confuse unified memory with dedicated VRAM

Apple Silicon is designed so CPU and GPU use the same physical memory. A 24GB Mac does not have “24GB for the CPU plus 24GB for the GPU.” The OS, browser, development tools and model share a single capacity. MLX provides a path suited to this design; NVIDIA CUDA code does not run unchanged on a Mac GPU.

Capacity affects whether a workload fits, while memory bandwidth and GPU configuration also affect speed. Even with the same 32GB capacity, chip generation, device, cooling and other running applications change speed. Swapping to an SSD may compensate for insufficient capacity, but it does not guarantee low-latency operation.

The MLX documentation describes arrays in shared memory; the Transformers cache strategies distinguish cache choices with different memory and speed tradeoffs. The numerical budgets below are original calculations, not provider hardware guarantees.

Build up from the lower bound for weights

A simple weight-size estimate is parameter count × bits per parameter ÷ 8. In decimal GB, 7B is approximately 14GB in FP16 or 3.5GB at 4 bits; 14B is about 28GB or 7GB; 32B is about 64GB or 16GB. These are rounded theoretical values. They exclude quantization scales, unquantized layers, layout overhead and temporary buffers. Do not mix GiB and GB displays.

Required memory ≈ weights + KV cache + activations/working memory + runtime + OS/other applications. Selecting a 4-bit model file does not automatically make the KV cache 4-bit. If you use KV-cache quantization, verify support and quality effects separately.

Context and concurrency increase other components

A general decoder KV-cache estimate is 2 × layers × KV heads × head dimension × tokens × bytes per element × concurrent sequences. Actual requirements vary with GQA, sliding windows and cache implementation. A hypothetical configuration of 32 layers, 8 KV heads, head dimension 128, FP16, 4096 tokens and one sequence gives approximately 0.5GiB; at 16384 tokens it gives about 2GiB. These are neither measurements on a machine nor specifications for a particular model.

Running four conversations concurrently can share weights, but each conversation adds cache and working memory. Estimating from maximum context alone is insufficient. Begin with one sequence, approximately 512–1024 input tokens and 128 output tokens, then increase context and concurrency separately.

A practical starting order and its limits

For a generic 24GB scenario, begin measuring with 0.5B–3B models, then consider a 4-bit 7B or 8B model. A 4-bit 14B model may also be a candidate at short context, but verify it together with other applications and cache. A 32GB scenario more easily provides headroom for the same experiments and leaves more room to try 4-bit 14B. A 4-bit 32B model already has a weight-only lower bound around 16GB; comfortable operation cannot be guaranteed in either scenario.

This is an order for trials, not a compatibility guarantee. Initially reserve at least several GB for the OS and other applications, for example 6–8GB, then increase the reserve according to actual memory pressure. Training additionally needs activations, optimizer state and other memory. Successful 7B inference does not establish that full-weight training of 7B is possible.

Calculation example: theoretical weights-only lower bounds

Status: estimated. Rounded decimal GB values exclude KV cache, quantization metadata, working memory and the OS. The 24GB/32GB cases are generic planning scenarios.

Scale FP16 4-bit
7B About 14GB About 3.5GB
14B About 28GB About 7GB
32B About 64GB About 16GB

Follow the workflow

  1. Use GB and GiB consistently.
  2. Compute the lower bound for weights.
  3. Reserve memory for the OS and other applications.
  4. Begin measurements with context 512, output 128 and concurrency 1.
  5. Stop if memory pressure worsens, swap grows substantially or execution stalls for a long time.

Review quality

  • Estimate the size of 4-bit weights.
  • Choose a testing order and stopping conditions for generic 24GB and 32GB scenarios.
  • Have you retained working memory rather than assigning all remaining capacity to KV cache?

Diagnose failures

Caveat What to check
Do not rank speed or intelligence by parameter count alone. Compare the same workload and settings, recording latency and quality separately.
Do not confuse active parameter count in an MoE model with memory required for all its weights. Estimate storage for all weights, then budget for the KV cache and working memory.
Do not treat 32GB of Mac unified memory and 32GB of NVIDIA GPU VRAM as identical conditions. Record the hardware and memory architecture; account for the OS and applications sharing unified memory.

Exercise: Estimate memory for 24GB and 32GB Macs

Status: not run (not-run).

Calculate the weight-only lower bounds for 7B, 14B and 32B models at 4-bit precision, then compute the remaining capacity in 24GB and 32GB scenarios.

Deliverable: A budget table explicitly labeled estimated.

Completion check: Have you retained working memory rather than assigning all remaining capacity to KV cache?

Use the experiment worksheet to plan conditions and record observations. Leave unknown measurements as null; record formulas and assumptions for estimates.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Apple MLX documentationPublished: Unknown · Accessed: 2026-10-03
02
Transformers KV cache ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
03
Ollama FAQ ↗docs.ollama.comPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.