kumyu.Learn
← Kumyu home日本語 ↗

REGISTERED TAG

Local inference

Run a model locally and examine memory, context and latency.

Published Updated
Beginner

Quantization, KV cache, and runtimes for a local LLM that stays responsive

Separate model weights, working memory, and generation speed; choose MLX, llama.cpp, or vLLM and turn the choice into a reproducible local evaluation.

11 min↗
Published Updated
Beginner

Understand models, tokens and inference

Follow the transformation from text to numbers, then from numbers to the next token.

8 min↗
Published Updated
Beginner

Choosing MiniMax and local LLMs: separate open weights, APIs, and feasibility

A practical research and experiment guide for comparing MiniMax and other models without conflating public repositories, weights, APIs, and inference servers.

11 min↗
Published Updated
Beginner

Estimate memory for 24GB and 32GB Macs

Budget weights, KV cache, working memory and the operating system separately.

11 min↗
Published Updated
Beginner

Read Hugging Face pages and model files

Check provenance, format and license so the model choice can be reproduced.

9 min↗
Published Updated
Beginner

Run a small LLM locally

First observe input, output and resource use, before judging answer quality.

10 min↗
Published Updated
Intermediate → local serving evidence

Local-model runtime evidence: evaluate Strata without inheriting its claims

Read a dated local-runtime release as a configuration and evidence contract: code and weight rights, local serving, release assets, author measurements, and a bounded offline fixture.

11 min↗
Published Updated
Beginner

RAG: retrieve knowledge instead of packing it into weights

Separate retrieval from generation for questions that need changing facts and traceable sources.

9 min↗
Published Updated
Beginner

Connect local AI to images and 3D

Apply a shared view of computing resources while understanding different production pipelines.

9 min↗
Published Updated
Beginner

Record experiments and integrate AI into a service

Turn small experiments into reproducible decisions and controlled operation.

18 min↗
Published Updated
Intermediate → serving design

vLLM 0.29–0.30: treat a runner default as a serving-boundary change

A dated release note on testing vLLM runner, queue, and route authorization changes before adopting a new serving release.

8 min↗
Published Updated
Intermediate → reliability evaluation

llama.cpp b11377: make JSON-schema output a parser-specific contract

A prerelease fixes a Ling 3.0 response-format gap; adopt it only through a pinned, parser-specific schema fixture.

10 min↗
Published Updated
Intermediate → decision evaluation

llama.cpp /v1/systemone: use a local decision endpoint as a routing signal

A dated local GGUF decision endpoint returns typed probabilities; it narrows routing, but never grants authority.

9 min↗
Published Updated
Intermediate → GPU failure diagnosis

llama.cpp CUDA MoE: treat a temporary-buffer dimension as a test boundary

A narrowly scoped prerelease fix shows why an MoE CUDA fixture must pin tensor shape, backend, and failure state before it changes a deployment.

8 min↗
Published Updated
Intermediate

Video model research notes: Provision formats and production choices after 2025

Check Veo, Sora, Wan, and LTX in official documents and compare cloud provision, open weights, inference code, and safety measures on different axes.

14 min↗
Published Updated
Introductory → production and delivery decisions

Choose local generation or Suno, then inspect the export

Compare execution responsibility, editable outputs and download conditions; distinguish sample rate, codec, container and loudness.

12 min↗