References are bundles of constraints, not decorations
If you generate a first video from only the text “I want to make a short film,” the first shot will often look striking. In the second shot the protagonist’s hairstyle changes; in the third, the room’s window has moved to the other side. The problem is not only a weak model. The creator has not supplied the world’s invariants. Reference-first production creates reusable visual facts before moving from text to moving image.
Google's Veo page introduces features such as scene/person/object/style reference images, character consistency, start/end frames, and camera control. This has more implications than a functional table of individual services. Even if we increase the number of conditions that generators can accept, it is up to people to decide what to fix and what to change. When you only have one reference image, don't think that it guarantees the sameness, including the back side, clothes when walking, and changes in facial expressions.
- 1One-sentence intent
- 2world invariants
- 3reference pack
- 1Reference pack
- 2character, prop, and location sheets
- 3shot list
- 1Start conditions for each shot
- 2generated candidates
- 3selection
- 4edit timeline
Divide the plan into “things that will not change” and “things that will change”
First, if you have 30 seconds, decide on the main character, location, time of day, and emotional changes in one sentence. For example, ``On a night after the rain, a delivery man with a red umbrella enters a quiet coffee shop and looks at the steam-clouded window to find relief.'' Next, create a fixed table. The protagonist's age and ethnicity are not vague labels, but are broken down into hair length and color, jackets, shoes, belongings, scars, accessories, and color combinations. It's not enough for the location to be "Tokyo-style." Specify the counter material, window position, lighting color, outside road surface, and prohibition of logos on the screen.
Place only the action and camera for each cut on the table to be changed. These are facial expressions, hand position, wetness, umbrella angle, focal length, angle of view, movement, and length. This allows you to determine whether the problem you want to fix during generation is a fixed requirement or a variable requirement. There is no need to rewrite the camera description to change the color of clothes.
How to create a reference pack
Instead of just one hero image, save at least the following with the same naming convention: These include frontal, diagonal, and side views of the person, the whole body and upper chest, a close-up of the face, close-ups of clothing and props, the wide view of the location, the materials and lighting of the location, and color samples. You cannot control the profile from the front view alone. It is not possible to maintain the grain of the cup handles and the wood grain of the table just by looking at the overall view of the place. Images can be generated, photographed, or 3D rendered, but make sure you have the right to use them and the consent of the person in the photo.
On the character sheet, also write permitted differences'' and prohibited differences.'' The raindrops can increase or decrease, the color of the eyes is fixed, the number of buttons on the jacket is fixed, the watch on the right wrist is fixed, and so on. Rather than increasing ambiguous aesthetic instructions, increase the number of facts that can be inspected when continuous cuts fail. Even in models that can use references, the longer the distance between the reference and the output, the more likely the model will break down. Turning around, sitting down, changing hands, and contact with multiple people are confirmed first by generating short tests.
Make shots into independent sub-problems
Do not put a 30-second sequence into one long prompt. Divide it into eight to ten shots, and limit each shot to one narrative intention, one primary action, and one camera movement. A shot card records its ID, duration, start and end references, framing, intended focal length, subject position, action, sound, and acceptance criteria. For example, V03 is four seconds: a medium shot from outside the window with the compression of a 50 mm lens; the protagonist closes an umbrella from right to down; only rain is heard; the watch and red umbrella must remain consistent.
- 1V01: establish the street, 3 s
- 2V02: umbrella and feet, 3 s
- 3V03: entering the café, 4 s
- 1V04: steaming cup, 3 s
- 2V05: expression, 4 s
- 3V06: window reflection, 3 s
- 1Each shot: reference + fixed constraints + primary action + camera intent + acceptance criteria
The advantage of generating each shot independently is that failures can be localized. You don't have to throw away V01 even if V03's fingers break down. The drawback is the connection. In order to pass the light, the position of the person, line of sight, direction of travel, and sound reverberation to the next cut, the last frame of the previous cut and the editing cut point are used as "transfer assets". This is especially effective for products that allow you to specify start and end frames, but it does not guarantee success.
Generate, sort, edit, audio
Create multiple low-cost candidates from each card. Rather than looking for good candidates within a single video, select each cut. The passing criteria is 0 to 2 points in 7 items: identity, movement, camera, background, physics, text/logo, and connection. For example, if the score is less than 11 out of 14 points, it will be regenerated, and if the person identity is 0, it will be decided to fail regardless of the total score. This is a reproducible design exercise, not a completed production result.
When editing, first check the connection without making any noise. Next, put environmental sounds, sound effects, dialogue, and music on separate tracks. Although native speech generation is fast, it is necessary to evaluate the meaning of the dialogue, mouth shape, noise, and rights processing. Check with and without sound, as sound can hide cuts in the video. Finally, we inspect the output length, subtitles, history display, and publication site regulations.
Record of failure strengthens next reference
You cannot reproduce the problem by simply writing that the prompt has been improved. For a failed shot, record the version of the input reference, model/version, seed, settings, candidate number, observed failure, correction hypothesis, and then the condition to change only one thing. If the person changes, before adding words or phrases, ask if there is not enough profile reference, if the movement is too large, or if the cut should be divided. The value of reference-first is not in getting the answer right every time. The goal is to be able to carry over which conditions dominate the screen to the next cut.
Out-of-the-box shot breakdown
Shot cards are not a competition for writing skills. This is a column for exporting variables that can be compared later. Fill in the ID / length / changes in the story / starting frame assets / ending frame assets / fixed attributes of the subject / variable movements / angle of view / movement / light / sound / things to pass on to the next / failure conditions for each cut. For V05, write ``4 seconds, from tension to relief, close view of the cup to the face in the window, red umbrella and watch fixed on the right wrist, place the cup and look at the window, stand still equivalent to 50mm, blue window light, the sound of rain weakens, the gaze is to the right, the disappearance of the watch and the movement of the window are not acceptable''.
With this format, when the window changes, you can choose whether to modify the background reference, angle of view, or editing, rather than adding the person's appearance. The same goes for character bibles. The character's name is Mina, her black hair falls behind her ears, her wet red raincoat, the watch on her right wrist, and a translucent umbrella are fixed, and her facial expressions can vary from expressionless to a little relieved. Rather than a strict resemblance of a face, we fix the facts that can be examined on a screen, such as color, position, and objects.
- 1Shot ID
- 2Fixed assets
- 3One primary action
- 4Camera intent
- 5Exit frame
- 1Exit Frame
- 2Next Shot Entry Reference
- 3Inspect Editorial Connection
- 1Rejection Criteria
- 2Cause Hypothesis
- 3Change one condition and try again
With this small specification, even if the model or editor changes, the same world can be carried over from shot to shot.
The reference pack after approval has a change history. If you overwrite a reference with a different clothing color with the same name, you can't keep track of which cuts used the old version. Just by leaving the edition name, date, and reason for the change, you can safely identify the scope of revisions to the series.
Provenance and reference consent after 2024
Reference-first work became more important as product interfaces accepted images, start/end frames, and character controls. It also changes the trust boundary: a reference is not merely an aesthetic input; it can contain a person's likeness, copyrighted design, metadata, or client context. Keep a source ledger for every reference: creator or licence, permitted transformation, audience, and deletion date. A visually stable shot is still unusable if its reference cannot be published.
The September 2026 community thread on multiple cuts versus a single generation should be read only as a creator workflow question. Use it to add a review checkpoint: before each new cut, compare the reference sheet against the generated last frame and explicitly approve the identity, wardrobe, scene geometry, and intended change.
A current reference workflow is still a role-assignment problem
Runway's September 24 Fashion Workroom guide is a current vendor example of using references across a production flow. It does not grant permission to reuse any reference or prove cross-shot identity. Translate it into a role sheet: subject identity reference, wardrobe/material reference, location/light reference, camera reference, and a negative reference for prohibited marks. Mark who supplied each asset, its permitted use, and its expiry.
The September 23 product-video guide also treats a multi-shot result as a planned sequence. Before each shot, freeze the approved upstream reference version. After each shot, write only the intentional delta—pose, prop state, or camera position. If a shot needs a new reference, create a new version rather than overwriting the prior evidence.
MENTAL MODEL / SHOT DESIGN
Give each shot one job.
Total: 12 seconds. Define a reference image, a start state, one action, and an end state for each shot. Arrange the still images in this order before generation to find transitions that do not communicate your idea.
SOURCES
01YOUR NOTES