Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.
- 1Fixed task and conditions
- 2measured, estimated or not-run record
- 3quality, cost, latency and permission checks
- 4adoption decision
- 5rollback plan
Before you begin
Prerequisite: Connect local AI to images and 3D. Estimated study time: 50 minutes.
This chapter is an editorial guide. The exercise and example record are UNRUN (not-run). Unknown measurements remain null.
What you will learn
- Keep records that clearly distinguish measurements, estimates and unrun plans.
- Make release decisions across four dimensions: quality, cost, latency and permissions.
Keep comparison conditions consistent
A speed report with only a model name cannot be reproduced. Record the model ID and revision, quantization, runtime version, hardware, input token count, maximum output tokens, actual output tokens, context, concurrency, batch size, and whether the run was cold or warm. Time to first token (TTFT), prefill speed, decode speed, total time and peak memory measure different things.
Measure the same input set several times and inspect the variation. Do not mix initial model loading with warm runs. For frameworks that process GPU work asynchronously, check synchronization or the official measurement method. The time taken by a Python function call alone is not necessarily the time taken by the GPU work.
The Transformers cache documentation explains memory and latency tradeoffs; the Ollama FAQ documents runtime and network configuration considerations. The worksheet records actual observed conditions so a provider’s supported feature is never confused with this course’s measured operation.
Make three evidence states explicit
measured means a value was actually measured under the stated conditions. estimated means a value was calculated from formulas, prices and assumptions. not-run means the plan has not yet been tried. Empty measurements must be null, not zero: zero could be misread as zero cost or zero memory. The code examples in this course are not-run; the capacity calculations are estimated.
For costs, record the currency, retrieval date of the price, billing unit and included items. API token counts and GPU time are not directly interchangeable. They become useful comparison inputs only after actual throughput has been measured for the model and hardware. Also record the number of completed tasks that pass the quality criteria as a denominator.
Conditions for moving development inference into operation
Configure maximum input and output sizes, timeouts, concurrency limits and rate limits in the backend. Validate not only whether JSON parses but also its schema and values, and cap retries after failures. Before executing a tool call returned by a model, the application must check permissions and arguments.
Minimize confidential information in generation logs and define retention and access. Manage the base model, adapter, tokenizer, prompt and retrieval index versions together so a model version can be rolled back after a problem. Like a code change, a model replacement should pass a fixed evaluation and staged checks.
Define what completion means
A suitable final project is Japanese terminology search and evidence-backed answers for a fictional product. Compare a small model, public or self-authored documents, RAG, and an adapter if needed. Explain what the system can and cannot do. The goal is to reproduce failures and choose remedies, rather than merely to run a powerful model.
Stopping conditions extend beyond going over budget. If you find data with unclear rights, a critical error, a state that cannot be measured, or a change that cannot be reproduced, step back and narrow the scope. The decision to add AI to a service must consider the effect on users and recovery from failures, not only a score for the model alone.
Roles and input → process → output
| Role | Responsibility |
|---|---|
| Learner | Design the hypothesis, data and evaluation. |
| Model | Attempt the specified transformation. |
| Application | Enforce limits, validation and permissions. |
| Stage | What it contains |
|---|---|
| Input | A fixed evaluation, the complete model artifacts, a runtime, a budget and operating conditions. |
| Process | Measure → compare → add guardrails → check operation in stages. |
| Output | A reproducible experiment ledger, an adoption decision and rollback procedures. |
Workflow
- Save the baseline experiment conditions.
- Distinguish measurements, estimates and unrun plans.
- Set minimum quality requirements and stopping conditions.
- Measure latency and memory at each concurrency level.
- Check permissions, logs and rollback.
Exercise: compare a small model with and without RAG
Status: UNRUN (not-run). Compare the same fictional-product question-answering task under two conditions: the small model alone, and the model with RAG.
Deliverable: an adoption decision memo containing quality, speed, memory, cost and failure examples.
Completion check: leave unmeasured fields as null, and keep the conclusion within the scope of the use case that was actually tested.
Quality checklist
- Does the record clearly distinguish measurements, estimates and unrun plans?
- Is the release decision evaluated for quality, cost, latency and permissions?
- Are unmeasured fields
null, and is the conclusion limited to the tested use?
Pitfalls and failure diagnosis
Do not advertise service speed from only the fastest result of a single run. A model's answer is not authorization to perform an external action.
| Caveat | What to check |
|---|---|
| Advertising service speed using only one fastest result. | Repeat the fixed workload and report cold/warm conditions, latency distribution and sample count. |
| Treating the model's answer as permission for an external action. | Keep execution permissions in the application; inspect the target and require the appropriate authorization. |
Experiment protocol: evidence states
The following protocol is the complete local AI course record specification. The example values describe a plan, not a completed run. null means unknown or unmeasured; it does not mean zero, and it does not establish that a component is absent.
| Status | Meaning |
|---|---|
measured |
A value actually measured under the specified conditions. |
estimated |
A value estimated from formulas, assumptions or public specifications. |
not-run |
A plan that has not yet been executed. Unmeasured numerical values remain null. |
Required record fields and the UNRUN example
Every field in this table belongs in the record. A configured limit or a chosen model can be known before execution; usage, timing, memory and quality remain unknown until measured.
| Required field | What to record | Example plan value |
|---|---|---|
status |
Whether the record is measured, estimated or unrun. | not-run |
measuredAt |
When the measurement was made. | null |
modelId |
The exact model identifier. | Qwen/Qwen2.5-0.5B-Instruct |
modelRevision |
The exact base model revision. | null |
tokenizerRevision |
The exact tokenizer revision. | null |
adapterId |
The adapter identifier, if applicable. | null |
runtime |
The inference or training runtime. | transformers |
runtimeVersion |
The runtime's exact version. | null |
hardware |
The hardware used for the run. | null |
os |
The operating system. | null |
quantization |
The quantization configuration. | none |
inputTokens |
The actual number of input tokens. | null |
maxOutputTokens |
The configured output token limit. | 64 |
actualOutputTokens |
The actual number of output tokens. | null |
contextLimit |
The configured context limit in tokens. | 512 |
concurrency |
The configured number of concurrent requests. | 1 |
batchSize |
The configured batch size. | 1 |
coldOrWarm |
Whether the run includes cold loading or uses a warm model. | null |
ttftSeconds |
Measured time to first token, in seconds. | null |
decodeTokensPerSecond |
Measured decode throughput, in tokens per second. | null |
totalSeconds |
Measured total elapsed time, in seconds. | null |
peakMemoryGB |
Measured peak memory, in GB. | null |
cost |
The cost fields listed below. | Unrun; amount and units unknown. |
quality |
The quality fields listed below. | Test version and results unknown. |
notes |
Context, qualifications and explanations. | “Course example. It has not been executed or measured.” |
Cost record
Keep the billing unit next to the amount. Include an explanation of what is counted; an empty includes list records that no included items have yet been specified, not that operation is free.
Field inside cost |
What to record | Example plan value |
|---|---|---|
status |
The evidence state of the cost. | not-run |
amount |
The billed or estimated amount. | null |
currency |
The currency of the amount. | null |
unit |
The billing unit. | null |
gpuHours |
GPU hours used or estimated. | null |
inputTokens |
Input tokens included in the cost record. | null |
outputTokens |
Output tokens included in the cost record. | null |
includes |
Items included in the cost calculation. | Empty list ([]). |
Quality record
Field inside quality |
What to record | Example plan value |
|---|---|---|
testSetVersion |
The fixed evaluation set version. | null |
passed |
The number of evaluated cases that passed. | null |
total |
The total number of evaluated cases. | null |
criticalErrors |
The number of critical errors. | null |
The example record has one note: “Course example. It has not been executed or measured.” Preserve that qualification when copying the plan. Do not replace the unknown results with guessed values.
Comparison protocol
- Keep the input and evaluation set fixed.
- Separate cold runs from warm runs.
- Keep concurrency fixed within a comparison.
- Report the distribution across multiple repetitions.
- Show units and formulas.
- Keep confidential information out of logs.
Use the experiment worksheet to plan a run or record what actually happened. The planned settings above must remain distinct from measured evidence.
Experiment worksheet
Copy the following plan into a local JSON file. Fill in the model and package revisions before running; leave timing, memory, token usage and costs as null until observed. Save each repetition separately with its evaluation-set version and raw output. This page provides a record format; it does not save observations automatically.
contextLimit is the planned total input-plus-output token budget, including chat-template tokens. For this 512-token plan with an output cap of 64, keep the tokenized input at or below 448. The earlier CPU example permits up to 512 input tokens separately; do not silently reuse that upper bound with this smaller total budget.
{
"status": "not-run",
"measuredAt": null,
"modelId": "Qwen/Qwen2.5-0.5B-Instruct",
"modelRevision": null,
"tokenizerRevision": null,
"adapterId": null,
"runtime": "transformers",
"runtimeVersion": null,
"hardware": null,
"os": null,
"quantization": "none",
"inputTokens": null,
"maxOutputTokens": 64,
"actualOutputTokens": null,
"contextLimit": 512,
"concurrency": 1,
"batchSize": 1,
"coldOrWarm": null,
"ttftSeconds": null,
"decodeTokensPerSecond": null,
"totalSeconds": null,
"peakMemoryGB": null,
"cost": {
"status": "not-run",
"amount": null,
"currency": null,
"unit": null,
"gpuHours": null,
"inputTokens": null,
"outputTokens": null,
"includes": []
},
"quality": {
"testSetVersion": null,
"passed": null,
"total": null,
"criticalErrors": null
},
"notes": [
"Course example. It has not been executed or measured."
]
}Use the local experiment ledger
Open the browser-local experiment ledger. Start with a planned record. Unknown values remain blank, and an executed record requires saved output evidence. The ledger does not run models or upload records. You can keep using the worksheet above as a separate document.
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01