One good image can still be a coincidence
The most dangerous thing about image generation is to take the first good-looking image and call it finished. A product's name is one letter different, a hand pierces the product, a logo that should not be used appears in the background, or a person's consent is published without knowing. The more subtle errors are, the more likely they are to be discovered after the layout is completed. In this exercise, participants will create an ad that ``shows off a new hypothetical unsweetened tea product with a person reading it by a window.'' Because it does not use real brands, people, or products, you can focus solely on recording and evaluating trials.
- 1Intended use
- 2Correct facts table
- 3Reference pack
- 4Composition rough
- 1Candidate generations
- 2scoring
- 3inpainting/compositing
- 4typography
- 5publication audit
- 1Reason for rejection
- 2Change one condition
- 3Regenerate
First, distinguish between “what may be generated” and “what may not be generated”
Do not ask the generator to supply product names, volumes, ingredients, prices, campaign dates, or legal text. Add them later as verified text layers. Generation can supply morning window light, a wooden table, steam, the atmosphere of a person reading, and the surrounding negative space. For the product itself, composite an approved pack shot on a separate layer or use it as a reference when exact shape matters.
Decide first whether the purpose is vertical for SNS, horizontal for web banners, or printing. If you create a picture without determining the proportions, the person's hands and the product will be cut off during the final trimming. For a 1080 x 1350 vertical image, place the safe area for the heading, the product area, and the margin where the person's line of sight is directed as a rough rectangle. Instead of creating a high-quality image, first create a composition that is easy to approve.
Limit the role of references and prompts
References fix the color of the product, the clothing of the person, the temperature of the light, and the material of the desk. Prompts convey the scene and screen composition, such as Reading by the window,'' Soft backlight in the morning,'' Product left, character right,'' and Copy space at the top.'' Rather than accumulating countless negative words, elements that should be absolutely avoided should be avoided by reference selection, masking, and composition. OpenAI's Image API Announcement mentions the use of incorporating image generation and editing into production tools, but does not guarantee the accuracy of the characters and trademarks being output.
Don't make just one large size candidate from the beginning. 4 to 8 candidates are presented in a small size and scored for composition, product distinctiveness, naturalness of the person, lighting, margins, and unnecessary text/logos. If there are two evaluators, hide the generation settings and choose the same scale to avoid confusing preference with suitability. The selected image is then upscaled or regenerated and compared with the original candidate.
- 1Rough composition + references + scene description
- 2low-resolution candidates
- 1Candidates
- 2score composition, hands, product, margins, and rights risk
- 1Selected candidate
- 2local corrections
- 3compositing pack shot and text layers
Local correction and layer editing
If the hand looks unnatural, instead of recreating the person's entire body, create a mask around the hand and the product to correct it locally. If steam covers the product label, remove it with a local masked edit. Mask correction leaves shadows and reflections at the border, so check at 100% display to see if the front and rear pixels naturally connect. Product logos, product names, prices, and warnings should be included in vectors or high-resolution approved materials. This is not a process that hides weaknesses in the model. It is designed to allow for updates and legal verification.
When using local open weights, there is an advantage that a large number of generation candidates can be generated. The official announcement of Stable Diffusion 3.5 guides you through multiple model versions and Community Licenses, but before running it, check the current model-weight license, GPU used, third-party node, and output storage path. For cloud APIs, review your organization's data policies, model updates, cost caps, and rate limits. Although the methods are different, the management of score sheets and approved materials is the same.
Final audit
Check factual text, product images, person consent, copyright/trademarks, background incidental text, skin/hands/reflections, accessibility alt text, and provenance records before publishing. The model name alone does not determine whether or not a generated image can be used for advertising. It depends on the references used, terms of service, organizational brand standards, and publication policy.
Finally, summarize the input, reference version, model/settings, candidates, scoring, mask, final layer, and approver. You don't need to reuse the same sentences for your next campaign. The idea is to reuse decisions about what should be fixed to maintain product-like quality, and what should be left to generation to speed up exploration.
Divide the hour-long production into three judgments
The first 20 minutes are spent just deciding on the composition. Place rectangles of copies, people, and products on a white background and compare about six candidates. For the next 20 minutes, choose the person and the light. Check to see if the person's hands are covering the product, if the line of sight is not too far away from the copy, and if there is white space around the product. The last 20 minutes are local corrections and compositing. Place product pack shot, product name, price, notes, and CTA as decision layers.
The scoring table gives 0 to 2 points each for composition, product identification, naturalness of the person, lighting, white space, incorrect information/unnecessary text, and copyright risk. If the product identification or rights risk is 0, it will be rejected regardless of the average score. If the product is too small, expand the area with a composition rough sketch before increasing the resolution. If your hands cover the label, mask only your hands and fix it. Text in the background is blurred or replaced, and typos in product names are placed from the approved text instead of being corrected by generation.
- 120 min: Composing with rectangles
- 220 min: Grading people and light
- 120 min: Mask Correction and Decision Layer
- 2Audit Facts/Rights/Views
- 1Post Publish Correction Request
- 2Safely Re-Output from Original Layer and Record
This order is not just for speed. Leave the corrected areas, remaining areas, and intentionally created layer boundaries. You can respond to price changes, regional notes, and seasonal campaigns without having to start over with your images.
In addition to the final image, the approval request should include a short description of the purpose, publication page, finalized text, generated parts, combined parts, usage references, and any remaining judgment points. If the approver gives the OK based on the atmosphere alone, confirmation of the product name and rights will be omitted. By arranging layout plans that allow you to select small screen previews, actual size, and text, you can check visuals and facts separately.
When a request is made to make a person look natural, the client is asked whether it is their smile, gaze, posture, hands, skin, or light that is causing the discomfort. For expressions, go back to the face area, for posture, refer to the person or pose, for light, go back to the background/shadow, and so on. File names that include purpose, ratio, language, edition, and approval date will reduce the risk of publishing old images with the same name.
For problems found after publication, just replace the image and do not finish. Make a short note of which inputs, which layers, which approval stages you missed, and add them to the next scoring sheet. For example, if your CTA cannot be read on a small screen, make readability at the final display size, not just the margin points, an essential gate. By returning the revision history to the process, the next production can begin with improved specifications rather than experience.
If you leave a one-sentence explanation of your reason for selecting the image, the next person will be able to reproduce your intentions.
Release checks after the 2024–2026 editing shift
The workshop should now keep an edit trail as carefully as its prompt trail: source asset, permission status, mask or crop, generated candidate, manual retouch, text proof, and final export. The recent OpenAI image announcement demonstrates that product surfaces can change quickly; it does not replace a visual release review. Inspect generated text at final pixel size and retain the original asset and export hash for the selected image.
MENTAL MODEL / REASONING ORDER
Change one thing at a time.
Decide what the image communicates, its medium and dimensions, and the meaning you want to preserve.
SOURCES
01YOUR NOTES