LEARNING PATH
LOCAL LLM
Run models locally and understand memory and speed.
2026-10-04textbook
LOCAL LLM · Introductory to practical
Quantization, KV cache, and runtimes for a local LLM that stays responsive
Separate model weights, working memory, and generation speed; choose MLX, llama.cpp, or vLLM and turn the choice into a reproducible local evaluation.
11 min↗
2026-10-04textbook
LOCAL LLM · Introductory to practical
Choosing MiniMax and local LLMs: separate open weights, APIs, and feasibility
A practical research and experiment guide for comparing MiniMax and other models without conflating public repositories, weights, APIs, and inference servers.
11 min↗