REGISTERED TAG
Local inference
Run a model locally and examine memory, context and latency.
Quantization, KV cache, and runtimes for a local LLM that stays responsive
Separate model weights, working memory, and generation speed; choose MLX, llama.cpp, or vLLM and turn the choice into a reproducible local evaluation.
Understand models, tokens and inference
Follow the transformation from text to numbers, then from numbers to the next token.
Choosing MiniMax and local LLMs: separate open weights, APIs, and feasibility
A practical research and experiment guide for comparing MiniMax and other models without conflating public repositories, weights, APIs, and inference servers.
Estimate memory for 24GB and 32GB Macs
Budget weights, KV cache, working memory and the operating system separately.
Read Hugging Face pages and model files
Check provenance, format and license so the model choice can be reproduced.
Run a small LLM locally
First observe input, output and resource use, before judging answer quality.
Local-model runtime evidence: evaluate Strata without inheriting its claims
Read a dated local-runtime release as a configuration and evidence contract: code and weight rights, local serving, release assets, author measurements, and a bounded offline fixture.
RAG: retrieve knowledge instead of packing it into weights
Separate retrieval from generation for questions that need changing facts and traceable sources.
Connect local AI to images and 3D
Apply a shared view of computing resources while understanding different production pipelines.
Record experiments and integrate AI into a service
Turn small experiments into reproducible decisions and controlled operation.
vLLM 0.29–0.30: treat a runner default as a serving-boundary change
A dated release note on testing vLLM runner, queue, and route authorization changes before adopting a new serving release.
llama.cpp b11377: make JSON-schema output a parser-specific contract
A prerelease fixes a Ling 3.0 response-format gap; adopt it only through a pinned, parser-specific schema fixture.
llama.cpp /v1/systemone: use a local decision endpoint as a routing signal
A dated local GGUF decision endpoint returns typed probabilities; it narrows routing, but never grants authority.
llama.cpp CUDA MoE: treat a temporary-buffer dimension as a test boundary
A narrowly scoped prerelease fix shows why an MoE CUDA fixture must pin tensor shape, backend, and failure state before it changes a deployment.
Video model research notes: Provision formats and production choices after 2025
Check Veo, Sora, Wan, and LTX in official documents and compare cloud provision, open weights, inference code, and safety measures on different axes.
Choose local generation or Suno, then inspect the export
Compare execution responsibility, editable outputs and download conditions; distinguish sample rate, codec, container and loudness.