Separate the model, inputs, generation, editing and verification.

Suggested reading time: about 8 minutes. This is an editorial production guide. Generation, performance and pricing comparison experiments have not been run.

Read first: A map of creative deliverables and work

What you will learn

  • Explain the difference between a model and an application.
  • Choose between generation and editing.
  • Add checks for AI’s weak points.

A model and a production tool are different

A model is a trained system that produces outputs from inputs. An application or API is the interface for operations such as supplying reference images, choosing dimensions and retaining history. Results can change when the model or settings change, even within the same service. Record the service name, the model identifier available to you and the execution date.

Generation has more than one mechanism

One entry point for understanding a language model is that it predicts the next processing units, called tokens, from the surrounding context. Tokens are not the same as characters or words. For a diffusion image model, a useful introductory picture is a process that gradually turns noise into an image matching the conditions. However, image and video models do not all share the same internal architecture. Do not claim that an undisclosed model “must use diffusion” or “must use this number of steps.”

Give each input a distinct role

Text communicates meaning and constraints. A reference image communicates appearance or composition. A mask identifies the area to change. Condition images such as pose or depth maps communicate structure. A model may not accept all of these inputs. New generation is useful for exploring candidates; editing is useful for correcting an accepted result while retaining its useful parts. Additional training and adapters are separate mechanisms, rather than simply longer prompts.

Verify proposals against the specification

An attractive image can still get text, left and right, props, reflections or identity features wrong. Generated code also carries no guarantee that it runs correctly or meets performance requirements. Adding the word “accurate” to a prompt does not replace a separate inspection method. As automation increases, input limits, failure handling and a route to manual correction become more important.

Work and responsibilities

  • AI production operator: designs conditions and settings.
  • Editor: selects and refines candidates.
  • QA reviewer: finds differences from the specification.

From input to output

Input Workflow Output
The minimum requirements.; Reference assets you have permission to use.; A model and output format that support the task. 1. Start with a small output to check that the task can run.; 2. Save the text, references and settings separately.; 3. Compare candidates against the acceptance criteria.; 4. Change one condition at a time.; 5. Edit only the necessary part of the selected candidate and compare it with the original. Candidates and the reasons for selecting one.; A generation record retaining the execution conditions.

Quality checks

  • Check the terms for using reference assets.
  • Confirm that the actual model supports the required features.
  • Inspect both the generated assets and the implemented result.

Diagnose failures

Symptom Possible cause Correction
More instructions make the result less stable. Conditions conflict or their priorities are unclear. Sort requirements into required, preferred and unwanted.
The same seed does not reproduce the result. The model, version, environment or settings differ, or the process is nondeterministic. Record the execution environment as well as the seed, and do not promise exact reproduction.

Exercise: Separate the roles of text and references

Status: not run (not-run). This is a plan with completion criteria. No artifacts, duration, cost, performance or success rate have been recorded.

Divide one character task into three stages: text only, text with a reference, and editing the selected image. Before running it, keep predictions and measured results in separate fields.

Deliverable: An input design sheet.

Completion criterion: You distinguish the model’s actual capabilities from your expectations.

When you run it, separate predictions from results and record conditions and evidence using the experiment notebook.

Continue into practice

Choose a use case and deliverable in the themed template library (Catalog candidates retained for review), then use Turn visual direction into words and specifications to refine the instructions.

Technical references and evidence scope

Technical references were checked on 2026-10-03. They support the technical concepts. The exercises, workflows and evaluation criteria are editorial proposals, rather than experiments performed by the source authors. Recheck changing APIs, prices and publication rules when using them.

Read the production flow

  1. 1Text and permitted references
  2. 2Candidate generation
  3. 3Specification review
  4. 4Targeted editing
  5. 5Acceptance check
Consider the sequence and each role.

This is an original conceptual process diagram; arrows show the review sequence, not a claim about an undisclosed model architecture.

Offline experiment worksheet

Condition Record before execution Record after execution
Input Prompt and references with reuse permission Exact saved inputs
Environment Intended service/model/version Actual identifier and execution date
Result Acceptance criteria Output path and observed defects
Effort and cost Budget limit Duration, usage units, currency and billed cost; unknown until measured

No model run was performed for this migration. Empty result cells mean unmeasured, not zero.

MENTAL MODEL / REASONING ORDER

Change one thing at a time.

Intent

Decide what the image communicates, its medium and dimensions, and the meaning you want to preserve.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Official documentationHugging Face Diffusers Text-to-image ↗huggingface.coPublished: Unknown · Accessed: 2026-10-03
02
Official documentationOpenAI Image generation guide ↗developers.openai.comPublished: Unknown · Accessed: 2026-10-03
Saved in this browser only.