Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.

  1. 1Fixed task and conditions
  2. 2measured, estimated or not-run record
  3. 3quality, cost, latency and permission checks
  4. 4adoption decision
  5. 5rollback plan
Consider the sequence and each role.

Before you begin

Prerequisite: Connect local AI to images and 3D. Estimated study time: 50 minutes.

This chapter is an editorial guide. The exercise and example record are UNRUN (not-run). Unknown measurements remain null.

What you will learn

  • Keep records that clearly distinguish measurements, estimates and unrun plans.
  • Make release decisions across four dimensions: quality, cost, latency and permissions.

Keep comparison conditions consistent

A speed report with only a model name cannot be reproduced. Record the model ID and revision, quantization, runtime version, hardware, input token count, maximum output tokens, actual output tokens, context, concurrency, batch size, and whether the run was cold or warm. Time to first token (TTFT), prefill speed, decode speed, total time and peak memory measure different things.

Measure the same input set several times and inspect the variation. Do not mix initial model loading with warm runs. For frameworks that process GPU work asynchronously, check synchronization or the official measurement method. The time taken by a Python function call alone is not necessarily the time taken by the GPU work.

The Transformers cache documentation explains memory and latency tradeoffs; the Ollama FAQ documents runtime and network configuration considerations. The worksheet records actual observed conditions so a provider’s supported feature is never confused with this course’s measured operation.

Make three evidence states explicit

measured means a value was actually measured under the stated conditions. estimated means a value was calculated from formulas, prices and assumptions. not-run means the plan has not yet been tried. Empty measurements must be null, not zero: zero could be misread as zero cost or zero memory. The code examples in this course are not-run; the capacity calculations are estimated.

For costs, record the currency, retrieval date of the price, billing unit and included items. API token counts and GPU time are not directly interchangeable. They become useful comparison inputs only after actual throughput has been measured for the model and hardware. Also record the number of completed tasks that pass the quality criteria as a denominator.

Conditions for moving development inference into operation

Configure maximum input and output sizes, timeouts, concurrency limits and rate limits in the backend. Validate not only whether JSON parses but also its schema and values, and cap retries after failures. Before executing a tool call returned by a model, the application must check permissions and arguments.

Minimize confidential information in generation logs and define retention and access. Manage the base model, adapter, tokenizer, prompt and retrieval index versions together so a model version can be rolled back after a problem. Like a code change, a model replacement should pass a fixed evaluation and staged checks.

Define what completion means

A suitable final project is Japanese terminology search and evidence-backed answers for a fictional product. Compare a small model, public or self-authored documents, RAG, and an adapter if needed. Explain what the system can and cannot do. The goal is to reproduce failures and choose remedies, rather than merely to run a powerful model.

Stopping conditions extend beyond going over budget. If you find data with unclear rights, a critical error, a state that cannot be measured, or a change that cannot be reproduced, step back and narrow the scope. The decision to add AI to a service must consider the effect on users and recovery from failures, not only a score for the model alone.

Roles and input → process → output

Role Responsibility
Learner Design the hypothesis, data and evaluation.
Model Attempt the specified transformation.
Application Enforce limits, validation and permissions.
Stage What it contains
Input A fixed evaluation, the complete model artifacts, a runtime, a budget and operating conditions.
Process Measure → compare → add guardrails → check operation in stages.
Output A reproducible experiment ledger, an adoption decision and rollback procedures.

Workflow

  1. Save the baseline experiment conditions.
  2. Distinguish measurements, estimates and unrun plans.
  3. Set minimum quality requirements and stopping conditions.
  4. Measure latency and memory at each concurrency level.
  5. Check permissions, logs and rollback.

Exercise: compare a small model with and without RAG

Status: UNRUN (not-run). Compare the same fictional-product question-answering task under two conditions: the small model alone, and the model with RAG.

Deliverable: an adoption decision memo containing quality, speed, memory, cost and failure examples.

Completion check: leave unmeasured fields as null, and keep the conclusion within the scope of the use case that was actually tested.

Quality checklist

  • Does the record clearly distinguish measurements, estimates and unrun plans?
  • Is the release decision evaluated for quality, cost, latency and permissions?
  • Are unmeasured fields null, and is the conclusion limited to the tested use?

Pitfalls and failure diagnosis

Do not advertise service speed from only the fastest result of a single run. A model's answer is not authorization to perform an external action.

Caveat What to check
Advertising service speed using only one fastest result. Repeat the fixed workload and report cold/warm conditions, latency distribution and sample count.
Treating the model's answer as permission for an external action. Keep execution permissions in the application; inspect the target and require the appropriate authorization.

Experiment protocol: evidence states

The following protocol is the complete local AI course record specification. The example values describe a plan, not a completed run. null means unknown or unmeasured; it does not mean zero, and it does not establish that a component is absent.

Status Meaning
measured A value actually measured under the specified conditions.
estimated A value estimated from formulas, assumptions or public specifications.
not-run A plan that has not yet been executed. Unmeasured numerical values remain null.

Required record fields and the UNRUN example

Every field in this table belongs in the record. A configured limit or a chosen model can be known before execution; usage, timing, memory and quality remain unknown until measured.

Required field What to record Example plan value
status Whether the record is measured, estimated or unrun. not-run
measuredAt When the measurement was made. null
modelId The exact model identifier. Qwen/Qwen2.5-0.5B-Instruct
modelRevision The exact base model revision. null
tokenizerRevision The exact tokenizer revision. null
adapterId The adapter identifier, if applicable. null
runtime The inference or training runtime. transformers
runtimeVersion The runtime's exact version. null
hardware The hardware used for the run. null
os The operating system. null
quantization The quantization configuration. none
inputTokens The actual number of input tokens. null
maxOutputTokens The configured output token limit. 64
actualOutputTokens The actual number of output tokens. null
contextLimit The configured context limit in tokens. 512
concurrency The configured number of concurrent requests. 1
batchSize The configured batch size. 1
coldOrWarm Whether the run includes cold loading or uses a warm model. null
ttftSeconds Measured time to first token, in seconds. null
decodeTokensPerSecond Measured decode throughput, in tokens per second. null
totalSeconds Measured total elapsed time, in seconds. null
peakMemoryGB Measured peak memory, in GB. null
cost The cost fields listed below. Unrun; amount and units unknown.
quality The quality fields listed below. Test version and results unknown.
notes Context, qualifications and explanations. “Course example. It has not been executed or measured.”

Cost record

Keep the billing unit next to the amount. Include an explanation of what is counted; an empty includes list records that no included items have yet been specified, not that operation is free.

Field inside cost What to record Example plan value
status The evidence state of the cost. not-run
amount The billed or estimated amount. null
currency The currency of the amount. null
unit The billing unit. null
gpuHours GPU hours used or estimated. null
inputTokens Input tokens included in the cost record. null
outputTokens Output tokens included in the cost record. null
includes Items included in the cost calculation. Empty list ([]).

Quality record

Field inside quality What to record Example plan value
testSetVersion The fixed evaluation set version. null
passed The number of evaluated cases that passed. null
total The total number of evaluated cases. null
criticalErrors The number of critical errors. null

The example record has one note: “Course example. It has not been executed or measured.” Preserve that qualification when copying the plan. Do not replace the unknown results with guessed values.

Comparison protocol

  • Keep the input and evaluation set fixed.
  • Separate cold runs from warm runs.
  • Keep concurrency fixed within a comparison.
  • Report the distribution across multiple repetitions.
  • Show units and formulas.
  • Keep confidential information out of logs.

Use the experiment worksheet to plan a run or record what actually happened. The planned settings above must remain distinct from measured evidence.

Experiment worksheet

Copy the following plan into a local JSON file. Fill in the model and package revisions before running; leave timing, memory, token usage and costs as null until observed. Save each repetition separately with its evaluation-set version and raw output. This page provides a record format; it does not save observations automatically.

contextLimit is the planned total input-plus-output token budget, including chat-template tokens. For this 512-token plan with an output cap of 64, keep the tokenized input at or below 448. The earlier CPU example permits up to 512 input tokens separately; do not silently reuse that upper bound with this smaller total budget.

{
  "status": "not-run",
  "measuredAt": null,
  "modelId": "Qwen/Qwen2.5-0.5B-Instruct",
  "modelRevision": null,
  "tokenizerRevision": null,
  "adapterId": null,
  "runtime": "transformers",
  "runtimeVersion": null,
  "hardware": null,
  "os": null,
  "quantization": "none",
  "inputTokens": null,
  "maxOutputTokens": 64,
  "actualOutputTokens": null,
  "contextLimit": 512,
  "concurrency": 1,
  "batchSize": 1,
  "coldOrWarm": null,
  "ttftSeconds": null,
  "decodeTokensPerSecond": null,
  "totalSeconds": null,
  "peakMemoryGB": null,
  "cost": {
    "status": "not-run",
    "amount": null,
    "currency": null,
    "unit": null,
    "gpuHours": null,
    "inputTokens": null,
    "outputTokens": null,
    "includes": []
  },
  "quality": {
    "testSetVersion": null,
    "passed": null,
    "total": null,
    "criticalErrors": null
  },
  "notes": [
    "Course example. It has not been executed or measured."
  ]
}

Use the local experiment ledger

Open the browser-local experiment ledger. Start with a planned record. Unknown values remain blank, and an executed record requires saved output evidence. The ledger does not run models or upload records. You can keep using the worksheet above as a separate document.

MENTAL MODEL / MEMORY

Separate model weights from KV cache.

Weights and KV cache grow independently. These values are planning estimates.

4.5 GBWeights 4.0 GB + KV 0.5 GB

GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Transformers KV cache ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
MLX LM official repositoryPublished: Unknown · Accessed: 2026-10-03
03
Ollama FAQ ↗docs.ollama.comPublished: Unknown · Accessed: 2026-10-03
04
Hugging Face Inference Endpoints ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.