The job of creating the same person is not the job of having to draw the face over and over again.

When you generate a person for a series, they may look attractive in the first photo, but the moment you change their pose, they become a completely different person. The more adjectives you add to this, considering it a weak prompt, the more likely it is that the composition will change along with the attributes you want to correct. To stabilize production, separate what should be handled by text, references, composition control, local editing, and learning.

References carry visual facts. Control carries positional relationships. Inpainting changes only part of it. LoRA is additive learning to bring the same features to many generations. Similar but not a replacement. If you just want to change a person's hair color, it's faster and requires less data management to try reference and mask editing before learning LoRA.

  1. 1Character Specifications
  2. 2Front/Landscape/Full Body Reference Set
  3. 3Generate by Purpose
  1. 1Composition Required
  2. 2Pose/Depth/Edge control
  1. 1Partial error
  2. 2mask
  3. 3inpainting
  4. 4retest
  1. 1Repeated Heavy Use
  2. 2Agreed Data
  3. 3LoRA Review
Consider the sequence and each role.

Create a reference pack as a shooting script

For people, collect frontal, 45-degree, profile, whole body, upper chest, expressionless face, smiling face, clothes, accessories, and hands. For a product, collect the front, side, back, material, shape without logo, and the situation in which it will be used. Write down what is confirmed information for each image. The red jacket is fixed, the background wallpaper is variable, and the mole on the left cheek is fixed. Including mixed lighting or different costumes as "correct" answers for the same person creates contradictions within the reference.

OpenAI's image-prompting document describes reference image/mask editing and settings to select high input fidelity. This is a means to preserve input detail and is not a guarantee that invisible angles will be reconstructed correctly. If the output differs from the reference, consider increasing the number of references, using a reference closer to the desired angle of view, reducing the change, and dividing the cut into two stages.

Control is stronger than “ask for composition”

If the person's arm or camera position is important, it will be easier to inspect if you provide the pose skeleton, depth, outline, and rough drawing rather than just saying, ``Raise your right arm.'' Pose conveys the position of joints, Depth conveys the context, and contours such as Canny convey the boundaries of the shape. Even so, the face, the material of the clothes, and even the fine text are not fixed. The example in which the official LTX-Video repository guides you to control Pose, Depth, and Canny shows that control extends to video, but the compatibility of the weights, UI, and versions to be used should be confirmed in each execution environment.

  1. 1Composition rough
  2. 2control image
  3. 3image suggestions
  1. 1Reference Image
  2. 2Identity Condition
  3. 3Image Suggestion
  1. 1Suggested
  2. 2Topical correction with mask
  3. 3Text/logo with layer
  4. 4Approved version
Consider the sequence and each role.

Inpainting closes modifications locally

It is risky to regenerate the entire image when you want to fix only the hand, the background sign, or the shadow. Specify the area to be changed with a mask and protect the area to remain. Make the mask border a little wider than the subject, and leave room for elements that cross the border, like hair or shadows, to naturally connect. After correction, a difference comparison is made to see if the face, clothes, logo, and text outside the mask have changed.

For text, correct product model numbers, and legal notations, use a design tool to add a final layer rather than having to use Inpainting over and over again. Generation is used for image foundation and exploration, and verifiable information is placed in a definitive way. This division of labor does not narrow expression. Leave freedom to safely modify in later revisions.

Decisions before using LoRA

LoRA is used as a method to adjust the weight of a part of an existing model as a small additional parameter and call out the characteristics of a specific person, costume, style, or product. The advantage is that you may be able to see consistent trends without sending lots of references each time. Disadvantages include small and biased data, including posture and background, excessive reproduction of similar images, compatibility with base model updates, data rights and consent, and usage conditions at the time of distribution.

The criteria for considering learning are whether the same subject will be used dozens of times, whether the passing rate is insufficient just by referencing, whether there is a right to use the training images and a path to deletion, and whether an evaluation set can be prepared in advance. First, separate 10 to 20 retained images for evaluation only. Even if we confirm that the images used for learning are similar, we cannot tell whether it is overfitting or generalization. We compare the base model, reference, and LoRA under the same conditions in a fixed test that includes faces, clothes, profiles, hands, different lights, and different backgrounds.

Definition of done

Completion does not mean ``the model was accurate.'' A record of the original reference, generation settings, mask, reason for adoption, editing layer, and usage rights remains, allowing another person to create the next piece in the same series. The more I work with people, products, and brands, the more identity emerges from the management of materials and approval rather than the performance of the model.

Minimum example of character bible

A person's identity cannot be tested if it is written as ``young woman, short hair, urbane.'' For the fictional character Yui, her features include shoulder-length black hair, a silver earring in her left ear, an unbleached shirt, a dark blue jacket, a thin ring on her right hand, and a blue-gray backpack. Variables include facial expression, posture, background, time of day, and wrinkles on the jacket. Prohibitions include changing hair length, moving earrings to the side, adding logos, and adding extra fingers. Corresponds to close-ups of the front, 45 degrees, side, whole body, hands, and bag.

LoRA does not mix training and evaluation. 40 approved photos will be used for training, and 12 photos that specify different lighting, posture, and background will be put on hold. Compare the basic model alone, reference only, and LoRA only, both with the same proportions and composition, and score not only the face but also clothes, bags, profile, and hands. Being able to learn is different from being able to distribute learned weights. Check the consent, deletion route, storage location, and origin of the material first.

  1. 1Person Specifications
  2. 2Approval Reference by Angle
  3. 3Create Criteria with Reference Generation
  1. 1Review Frequency and Consent
  2. 2LoRA Candidates
  3. 3Compare on Hold Evaluation Sets
  1. 1Discrepancies
  2. 2Separate Data/Settings/Usage
  3. 3Return to Adoption or Reference Steps
Consider the sequence and each role.

After Inpainting, don't just look at the areas you fixed. Check whether the skin color, shadows, seams of clothes, and background perspective outside the mask are continuous with the original image at 100% display and at the final publication size. Local corrections are small operations, but they are also operations that determine the reliability of the series.

When the series exceeds 10 pages, arrange the final versions on a contact sheet. If you look at hair color, main color of clothes, position of accessories, and product proportions, you will notice drifts that you might have missed if you checked each photo one by one. When discovered, we will update future standard references before re-creating all past works, and separately judge the necessity and cost of correction for published assets. When using a portrait of a person, separate access for reference, LoRA, and output, and do not retain personal data for more than necessary.

In the prompt to be reused, write not only the person's unique name, but also the reference version to be used with. This is because the same sentence is not necessarily interpreted with the same meaning when the model is changed. For month-to-month comparisons, use fixed prompts, fixed references, and fixed rating tables to distinguish whether changes are model- or material-based. Reproducibility comes not from fixing seeds once, but from leaving conditions and judgments in place.

Community signal: iteration cost is a workflow concern

A September community roundup mentioned local image tooling and editing nodes. It is a discovery signal, not a benchmark or a safety claim. Use it to ask whether a proposed LoRA or node reduces repeat work on the same controlled reference set. Version the base model, adapter, seed, mask, and compositing step; otherwise an apparent consistency improvement cannot be reproduced or attributed.

As a dated 2024 reference, Stable Diffusion 3.5 is relevant here because it separates a model announcement from the later adapter, reference, and inpainting workflow that must be versioned independently. It is not a current recommendation.

MENTAL MODEL / REASONING ORDER

Change one thing at a time.

Intent

Decide what the image communicates, its medium and dimensions, and the meaning you want to preserve.

SOURCES

01
GPT Image 1 reference ↗developers.openai.com · unknown
02
LTX-Video official repository ↗github.com · unknown
03
Introducing ChatGPT Images 2.5 ↗openai.com · 2026-09-08
04
StableDiffusion community roundup (community experience) ↗www.reddit.com · 2026-10-01
05
Introducing Stable Diffusion 3.5 ↗stability.ai · 2024-10-22

YOUR NOTES