Keep tokens, images, seconds and credits distinct
Suggested reading time: about 8 minutes. This is an editorial production guide. Generation, performance and pricing comparison experiments have not been run.
Learning goals
- Separate usage, billed amounts and quality.
- Record comparison experiments reproducibly.
- Distinguish illustrations from measurements.
Keep measurement separate from interpretation
- 1Hypothesis and fixed conditions
- 2Attempts including failures
- 3Outputs and usage evidence
- 4Quality, latency and human time
- 5Cost with units and date
- 6Scoped conclusion
- 1Missing measurement
- 2Keep null and its status
- 3State the evidence gap
This original evidence diagram separates collection from interpretation. Missing values stay visible through the conclusion; when no output is accepted, cost per accepted output has no defined value.
Who does this work?
- Production artist: record attempts and acceptance decisions.
- Developer: measure usage and performance.
- Producer: judge total time and cost.
Inputs
- A hypothesis to compare.
- Conditions to hold fixed and conditions to change.
- Model, settings and date.
- Available usage records and pricing information.
Production workflow
- Define evaluation criteria that match the use case before running.
- Record a baseline under the same input conditions.
- Change one variable and make multiple attempts.
- Record usage, waiting time, editing time and acceptance decisions.
- Choose the next change using both quality and total cost.
Outputs
- An experiment-record JSON file.
- References to outputs and logs.
- Limitations and the next hypothesis.
Tokens are not a universal currency
Text input and output, image input and generated images, audio and video can have different usage measures and prices even within the same service. Do not convert every image to a fixed token count. Seconds of generated video, seconds spent waiting for generation and time renting a GPU are distinct quantities. Credits are a service-specific unit; without their conversion rules, they cannot be compared directly between services.
Collecting cost evidence
When possible, save API response usage, billing records and execution logs. Label unavailable usage as not collected rather than setting it to zero. Match pricing to model, settings, date and currency. Keep a subscription’s monthly fee distinct from the marginal cost of a single run, and state the allocation rule if you divide a subscription cost between runs. Using a free allowance does not mean no computing resources were consumed.
Look at cost per accepted output
If you generate ten candidates and accept one, the production cost of that accepted image includes every candidate and the editing time. Even a low API price may produce a high overall production cost when extensive shape corrections are needed. Recording quality scores, acceptance rate, time and charges separately helps you judge what suits your own use case.
One experiment is not enough to generalize
Run multiple attempts on the same subject and change one condition at a time. Model updates, device, screen resolution and background processes can affect results. Keep failures alongside successful examples and state when the sample is small. Record the evaluator and criteria for subjective scores, and distinguish them from objective performance measurements.
How this course labels evidence
Measured means a result actually obtained with logs or outputs; illustrative means hypothetical values used to explain a mechanism; not-run means only a plan exists. Not-collected means a run occurred but that particular value was not obtained. Display these with different labels and presentation. The exercises in this course have not been run. A prompt example neither guarantees a particular quality nor indicates that a generation test has been completed.
Quality checklist
- Unknown values have not been replaced by zero.
- The model and pricing reference dates are clear.
- Failures and retries are included in cost.
- Measurements, illustrations and predictions are explicitly labeled.
- Public logs contain no confidential images or secrets.
Diagnose failures
| Symptom | Likely cause | Next step |
|---|---|---|
| A cheaper model costs more overall. | Retries and corrections were not counted. | Compare total cost per accepted output and human working time. |
| FPS differs for the same demo. | Device, resolution, application state or other conditions differ. | Align the conditions and retain the distribution of frame times as well. |
Exercise: Compare two production methods
Status: not run (not-run). This is a planned exercise, not a measured result.
Task: Plan to create the same use case with methods A and B. Define quality criteria, maximum budget and stopping conditions before running, and never display an unexecuted row as measured.
Deliverable: A comparison plan and experiment record.
Completion criteria: Explain cost units, provenance, acceptance rate and values that were not collected.
First check what is being counted
No provider prices are fixed in this course. Retrieve current official pricing for the exact model, settings, date and currency.
| Usage unit | Record | Interpretation check |
|---|---|---|
Text input and output tokens (text_tokens) |
Separate input, output, cache and other categories exactly as the API defines them. | Do not double-count cached tokens if the API includes them in total input. |
Image input and output tokens (image_tokens) |
Save image-specific usage when available. | Do not apply a universal conversion from pixels or image count. |
Number of generated images (image_count) |
Size, quality, candidate count and retries. | Check whether the provider charges per image or per token. |
Duration of generated video or audio (output_seconds) |
Seconds, resolution, model and whether audio is included. | This is not generation waiting time. Units and settings depend on the provider. |
Computing resource time (compute_seconds) |
GPU type, seconds used and whether idle time is included. | Do not apply an API token price to local execution without a basis. |
Service credits (credits) |
The before-and-after balance difference and the date the conversion rules were obtained. | Units with the same name can have different values between services. |
Human editing time (human_minutes) |
Time spent selecting, editing, implementing and reviewing. | If you assign an hourly rate to your time, state that it is an assumption. |
Total estimated API cost = Σ (usage for each nonoverlapping billed item ÷ the price’s base unit × its unit price). Cost per accepted output = total including retries ÷ accepted output count. Display this as undefined when the accepted count is zero.
Illustrative arithmetic only
A fictional example for learning the arithmetic. These are not the prices of a real service. (illustrative)
Assumptions
| Field | Recorded value |
|---|---|
| Generated images | 6 |
| Accepted images | 2 |
| Fictional price per image (JPY) | 10 |
| Human editing time (minutes) | 20 |
6 images × a fictional ¥10 = ¥60. With 2 accepted images, the cost is ¥30 per accepted image. Record the 20 minutes of human editing separately.
Excluded from this arithmetic
- Tax
- Allocation of subscription charges
- Equipment costs
- Electricity
- Labor costs
A reproducible experiment protocol
- Write a hypothesis and acceptance conditions suited to the use case before running.
- Record inputs, model, settings, references and device.
- Change one condition and keep the comparison baseline.
- Save every attempt and failure.
- Evaluate quality, waiting time, cost and human time separately.
- State the scope of the conclusion and the next hypothesis.
Required provenance
- Execution date and time.
- Operator or environment.
- Model identifier and any available version.
- Provenance and usage permission for inputs and reference assets.
- Settings and whether a seed is supported.
- References to outputs or logs.
- Source and units of usage records.
- Pricing URL and date retrieved.
- Evaluation criteria and evaluator.
Record fields and value types
This field reference describes the full creative record format. It is a specification, not a completed run. Preserve identifiers exactly when exporting records, and use null for unknown measurements.
| Field | Expected value |
|---|---|
id |
string |
status |
not-run | measured | illustrative |
hypothesis |
string |
taskType |
image | image-edit | 3d | realtime | video | workflow |
runAt |
ISO 8601 | null |
provider |
string | null |
modelId |
string | null |
modelVersion |
string | null |
input.prompt |
string |
input.referenceAssets |
array of {id,source,permission,hash?} |
input.parameters |
object |
input.seed |
number | string | null |
input.seedSupported |
boolean | null |
environment.appVersion |
string | null |
environment.device |
string | null |
environment.os |
string | null |
environment.browser |
string | null |
environment.renderResolution |
string | null |
environment.gitCommit |
string | null |
outputs |
array of {assetId,pathOrUrl,hash?} |
usage |
array of {kind,quantity:number|null,unit,source,status:measured|not-collected|not-run} |
cost.amount |
number | null |
cost.currency |
string | null |
cost.kind |
billed | estimated | illustrative | not-collected |
cost.priceSourceUrl |
string | null |
cost.priceCheckedAt |
date | null |
cost.formula |
string | null |
timing.latencySeconds |
number | null |
timing.humanEditingMinutes |
number | null |
timing.sampleDurationSeconds |
number | null |
quality.rubric |
array of {criterion,scale} |
quality.scores |
array of {criterion,value,evaluator} |
quality.accepted |
boolean | null |
quality.failureNotes |
string[] |
performance.medianFrameMs |
number | null |
performance.p95FrameMs |
number | null |
performance.measurementMethod |
string | null |
conclusion |
string | null |
limitations |
string[] |
A comparison plan that has not been run
The following plan has no measured outputs, usage, price, timing, quality scores or frame-time results. Empty collections and null values describe missing evidence; they do not represent zero resource use.
| Field | Recorded value |
|---|---|
id |
ocean-wave-comparison-plan |
status |
not-run |
hypothesis |
Can separating small waves into normals add appearance more cheaply than increasing the vertex count? |
taskType |
realtime |
runAt |
Unknown / not collected (null) |
provider |
Unknown / not collected (null) |
modelId |
Unknown / not collected (null) |
modelVersion |
Unknown / not collected (null) |
input.prompt |
Compare vertex-only waves with waves combined with normals, using the same camera and screen size. |
input.referenceAssets |
No entries yet (not run) |
input.parameters |
No entries yet (not run) |
input.seed |
Unknown / not collected (null) |
input.seedSupported |
Unknown / not collected (null) |
environment.appVersion |
Unknown / not collected (null) |
environment.device |
Unknown / not collected (null) |
environment.os |
Unknown / not collected (null) |
environment.browser |
Unknown / not collected (null) |
environment.renderResolution |
Unknown / not collected (null) |
environment.gitCommit |
Unknown / not collected (null) |
outputs |
No entries yet (not run) |
usage |
No entries yet (not run) |
cost.amount |
Unknown / not collected (null) |
cost.currency |
Unknown / not collected (null) |
cost.kind |
not-collected |
cost.priceSourceUrl |
Unknown / not collected (null) |
cost.priceCheckedAt |
Unknown / not collected (null) |
cost.formula |
Unknown / not collected (null) |
timing.latencySeconds |
Unknown / not collected (null) |
timing.humanEditingMinutes |
Unknown / not collected (null) |
timing.sampleDurationSeconds |
Unknown / not collected (null) |
quality.rubric |
criterion: Agreement between the wave silhouette and lighting. / scale: Subjective score from 1 to 5; record the evaluator. |
quality.scores |
No entries yet (not run) |
quality.accepted |
Unknown / not collected (null) |
quality.failureNotes |
No entries yet (not run) |
performance.medianFrameMs |
Unknown / not collected (null) |
performance.p95FrameMs |
Unknown / not collected (null) |
performance.measurementMethod |
Unknown / not collected (null) |
conclusion |
Unknown / not collected (null) |
limitations |
Performance differences are unconfirmed because the plan has not been run.; Results depend on device, resolution and shader implementation. |
Experiment worksheet
Copy this example into a local JSON file when planning a run. It stores no data on this site and does not start a model call. Keep the schema identifiers unchanged. Empty arrays and null values mean missing evidence, not zero usage; add usage entries only with their source, unit and measured/not-collected/not-run status. Keep the original plan when adding the executed record.
{
"id": "creative-comparison-plan",
"status": "not-run",
"hypothesis": "Replace with the comparison question",
"taskType": "workflow",
"runAt": null,
"provider": null,
"modelId": null,
"modelVersion": null,
"input": {
"prompt": "",
"referenceAssets": [],
"parameters": {},
"seed": null,
"seedSupported": null
},
"environment": {
"appVersion": null,
"device": null,
"os": null,
"browser": null,
"renderResolution": null,
"gitCommit": null
},
"outputs": [],
"usage": [],
"cost": {
"amount": null,
"currency": null,
"kind": "not-collected",
"priceSourceUrl": null,
"priceCheckedAt": null,
"formula": null
},
"timing": {
"latencySeconds": null,
"humanEditingMinutes": null,
"sampleDurationSeconds": null
},
"quality": {
"rubric": [],
"scores": [],
"accepted": null,
"failureNotes": []
},
"performance": {
"medianFrameMs": null,
"p95FrameMs": null,
"measurementMethod": null
},
"conclusion": null,
"limitations": [
"Comparison has not been run; no result is available."
]
}Related Catalog candidates and practice
Prerequisite chapters: Understand AI generation and its limits
The browser-local ocean shader lab uses Three.js 0.186.1 with WebGLRenderer and a local vendored module; it makes no model API calls. Its stages are a flat plane, vertex waves, and approximate light and color. It does not implement physical ocean simulation, true reflection or refraction, or shoreline foam. Opening it does not complete the comparison exercise or establish performance on other devices.
Open the procedural ocean shader lab
Use the experiment worksheet for your plan and evidence
Primary sources
- OpenAI API Pricing — verified: 2026-10-03
- OpenAI Image generation guide — verified: 2026-10-03
- MDN requestAnimationFrame — verified: 2026-10-03
Source scope checked on 2026-10-04
The current pricing page distinguishes models and modalities. This check does not supply a measured bill or freeze future prices. Copy the exact rate, billing unit and retrieval date into the worksheet only when designing or executing an authorized comparison.
The linked ocean lab is a separate practice asset. Keep its version, device and actual stage names with any later measurements. The chapter’s planned exercises and capstones remain not run.
MENTAL MODEL / REASONING ORDER
Change one thing at a time.
Decide what the image communicates, its medium and dimensions, and the meaning you want to preserve.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01