“Public” describes four different things
When you encounter a name such as MiniMax, it is tempting to conclude that a GitHub repository means local inference is possible. Public source repositories, model weights, APIs, and hosted UIs are separate assets. The MiniMax GitHub organization contains public projects, but a repository README does not by itself grant access to all model weights or establish that a particular API model is reproducible. The official service surface is the platform documentation; confirm contracts, regions, and enabled features in the current account and documentation.
- 1Public code
- 2implementation can be read
- 3availability of weights needs separate confirmation
- 1Public weights
- 2candidate for local inference
- 3inspect license and format
- 1Cloud API
- 2callable capability
- 3inspect price, retention, and region
- 1Hosted UI
- 2path for trying a product
- 3does not imply an API or weights
This distinction is not unique to MiniMax. Even when names share a suffix, an API version may include tool calling, search, guardrails, and server-side updates. An open-weight version may be available, yet matching output still requires fixing the tokenizer, chat template, recommended runtime, long-context settings, and license terms yourself. The first row of a comparison sheet should state what was acquired, before it states performance.
Records that survive a changing model catalog
“Latest” decays quickly. The model list and regions visible on official pages when this textbook was created on 2026-10-04 can change later. Do not invent future model names or rankings. For every candidate, separately record the official model-card or release-note URL, announcement date, access date, weights URL, license, API model ID, and observed features. When an announcement date cannot be read from the source, record unknown; do not replace it with the access date.
The Hugging Face model-card guidance is a useful checklist for intended use, limitations, evaluation, and licensing. A model card can contain an inference command without implying equal speed on a Mac, Windows machine, and GPU cluster. Treat its figures as provider measurements unless an independent comparison says otherwise.
A comparison sheet needs at least:
- Acquisition form: API only, published weights, code only, or a combination.
- Rights: separate columns for weights, code, outputs, commercial use, and redistribution.
- Operations: offline capability, transmitted data, retention, region, and where keys live.
- Practical fit: Japanese, structured output, tool use, context, and streaming; leave unverified cells empty.
- Reproduction: model revision, quantization, chat template, runtime, hardware, and run date.
Why choose an API, and why choose weights?
An API lets you test capability without downloading a model or operating GPUs. Provider-side updates, safety controls, and extensions may improve over time. The trade-offs are pricing, network dependence, rate limits, data handling, and model updates outside your control. Before sending personal or non-public learning material, read the provider’s data policy and your organization’s handling rules. Never embed a key in a browser. Even a small relay server must avoid logging body text or credentials.
Weights let you run through network loss, control where input resides, and fix quantization and system prompts. The cost is storage, power, vulnerability response, an execution environment, and quality verification. “Local” does not mean free: hardware, time, operations, and possibly commercial licensing all cost money. Instead of declaring an API or local model superior, draw boundaries around the job: drafts, internal retrieval, batch extraction, or low-latency dialogue.
Compare failure shapes, not one impressive response
Even if candidate A excels on a math benchmark and candidate B excels at code, neither result proves suitability for your grading task or course notes. Start with 10–30 small representative tasks: summaries that preserve proper nouns, questions that must refuse unsupported guesses, JSON under a specified schema, short code fixes, and citation extraction from long documents. Record unsupported assertions, malformed JSON, over-obedience to instructions inside input, waiting time, and cost alongside correct answers.
- 1Work failure examples
- 220 representative tasks
- 3API and local candidates
- 1Same prompts and conditions
- 2automated checks plus human review
- 3write adoption criteria
- 1Model update
- 2rerun the fixed set
- 3record the difference
Structured output is an especially common trap. If a provider implements JSON Schema or tool calling as a server feature, raw local weights in the same model family may not have that constraint. If you add grammar constraints, retries, or validators locally, include that code in the comparison. A “model alone” and an API product must not occupy one undifferentiated row, or you cannot tell what reduced failures.
Practice: a research card that includes MiniMax
- Open the official GitHub organization and platform documentation, then save primary URLs for the capability you need—text, speech, image, or video. Use third-party summaries only to discover candidates.
- For every candidate, fill four URL fields: direct weights distribution, model card, license text, and API reference. Mark a missing field as unverified.
- Before a billable API call, read the request schema, error types, pricing, and data policy. Do not infer availability from these documents.
- Only for candidates whose weights are explicitly published, verify that your local runtime supports the format. Do not claim a conversion preserves quality simply because it completes.
- Build a small evaluation set and run one candidate at a time later, in an environment where a key can be configured. Label results as unexecuted, run on this machine, or provider evaluation.
Selecting a model is not selecting the largest number. It decides where data can reside, whether the same output can be reproduced, and whether failure can be detected. Names may change; these card fields remain, so the decision footing survives a fast-moving field.
Availability is a moving contract
The 2024–2026 local-model cycle repeatedly shows why a provider name cannot answer “can I run this locally?” An August 2026 LocalLLaMA thread raises model-family and KV-cache questions; it is community discussion, not confirmation of MiniMax weights, licence, or performance. For each MiniMax candidate, obtain the named model's official card or repository, licence, checksum/format, supported runtime, API documentation, and account-specific terms on the same day. If any link is missing, record “unverified” and do not substitute a similarly named model.
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
SOURCES
01YOUR NOTES