Original conceptual diagram. Arrows show the dependency or decision sequence, not measured performance.
- 1Rights and output format
- 2small generation trial
- 3production correction
- 4mesh or image checks
- 5delivery in target tool
Before you begin
Prerequisite: Handle medical terminology: sources, negation, units and human review. Estimated study time: 50 minutes.
This chapter is an editorial guide. The exercise is UNRUN (not-run). References to 24 GB and 32 GB machines describe general memory configurations, not measured performance or a particular person's hardware.
What you will learn
- Explain why the settings that dominate resource use differ between LLMs and image generation.
- Separate validation of a 3D deliverable from evaluation of a generative model.
Resolution and concurrency also matter for images
Image generation can involve several components beyond text conditioning, such as an image encoder, denoiser and decoder. Consider resolution, batch size, generation steps, precision, auxiliary models and upscalers when estimating resource needs. Instead of comparing only LLM tokens per second, record seconds per image together with the settings and peak memory.
Diffusers has an Apple MPS path, but support for the specific pipeline and operations must be checked. The ability to perform some image generation on a 24 GB or 32 GB machine does not establish that a large video model or image training will run comfortably. Begin validation with a small resolution, batch size 1 and a single pipeline.
The Diffusers MPS documentation provides a specific Apple Silicon image-pipeline path and notes batch-inference limitations. It supports checking the pipeline individually, rather than inferring general video or 3D support from an LLM result.
3D has work after generation
The outputs of text-to-3D or image-to-3D models vary by method: meshes, textures and point clouds are different deliverables. A pleasing rendered image does not establish that a mesh is closed, that its polygon count is suitable, that UVs and materials are well organized, or that it can support a rig and animation.
Separate generation, correction, retopology, texture adjustments, export and display inside an engine. Also distinguish the data needed to train a 3D machine-learning model from the work of editing a finished 3D asset in production software. Rather than forcing a CUDA-only research implementation to run on a Mac, moving only the necessary stage to a GPU platform may sometimes be the practical choice.
Continue into the creative production course
The local AI course teaches a common structure: model and input, computing, output and evaluation. The creative production course explains how to finish those outputs as images, videos or 3D works. Use the same experiment ledger, keeping unrun expectations, calculated estimates and measured values separate.
Image and 3D inputs also raise questions about usage rights, a person's consent and the permitted publication scope. Check the ability to generate locally separately from permission to distribute or use the result commercially. Define completion not only by processing speed but also by whether the asset can be edited for the intended purpose and used after export.
Read the AI creative production course to continue into image and 3D production and finishing.
Roles and input → process → output
| Role | Responsibility |
|---|---|
| Learner | Design the hypothesis, data and evaluation. |
| Model | Attempt the specified transformation. |
| Application | Enforce limits, validation and permissions. |
| Stage | What it contains |
|---|---|
| Input | Prompts, images or 3D material with verified usage rights, plus settings. |
| Process | Generate → correct in production software → inspect for the intended use. |
| Output | An image or 3D asset, with the generation conditions and rights record. |
Workflow
- Decide the output format and where it will be used.
- List the required model components.
- Measure peak memory with small settings.
- Evaluate generation quality separately from usability in production.
- Continue to the finishing stages in the creative production course.
Exercise: compare generation with delivery
Status: UNRUN (not-run). For one image and one game-ready 3D asset, write three checks that are needed after generation for each.
Deliverable: a comparison between successful generation and successful delivery.
Completion check: a good-looking image alone must not be treated as proof that a 3D asset is complete.
Quality checklist
- Can you explain why LLMs and image generation have different dominant settings?
- Have you separated 3D deliverable validation from generative model evaluation?
- Are you checking the actual 3D asset rather than judging completion from a pleasing image alone?
Pitfalls and failure diagnosis
Do not transfer an LLM memory estimate directly to images, video or 3D. Do not guarantee unverified Mac support or processing speed.
| Caveat | What to check |
|---|---|
| Reusing an LLM memory estimate unchanged for image, video or 3D work. | Record resolution, batch size, generation steps, auxiliary models and peak memory for the target workflow. |
| Promising Mac compatibility or processing speed that has not been checked. | Check the specific pipeline and operations; begin with small settings and record the result on the actual device. |
Use the experiment worksheet to plan or record this experiment.
MENTAL MODEL / MEMORY
Separate model weights from KV cache.
Weights and KV cache grow independently. These values are planning estimates.
GB uses 10⁹ bytes. KV assumes 32 layers, 8 KV heads, head dimension 128, FP16, and batch size 1. Quantization metadata, runtime buffers, the OS, and model-specific structure need additional memory. For MoE, distinguish total from active parameters.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01