30 seconds is short. So it's long enough for verification

If you think you can make something in 30 seconds overnight, you will fail. The 30 seconds have an introduction, change, and landing, and consistency is needed because people and places are shown multiple times. Even so, it is more suitable than a long story because all the conditions can be contained in one record sheet. This chapter is not a record of the production of completed works. This is a design exercise that can be applied to any video generation service.

The premise is: “After closing at a bookstore, a clerk returns a book to the shelf and the rain outside the window stops.” Use one character, one location, one prop, and a small emotional shift. Avoid spectacular transformations and crowds, not to make the task model-friendly, but to make the source of a failure diagnosable.

  1. 130-second intention
  2. 28-cut card
  3. 3reference pack
  1. 1Candidate for Low Resolution
  2. 2Scoring
  3. 3High Quality with Passing Cut
  1. 1Edit Silence
  2. 2Sound and Subtitles
  3. 3Final Audit
Consider the sequence and each role.

Design made in the first 60 minutes

In the first 10 minutes, write the invariants. They are the clerk's hair, shirt, apron, glasses, book binding, the color of the shelves, the position of the window, and the blue light after the rain. In the next 20 minutes, prepare references for people, books, and stores. When referencing a product, record the version and terms of use. With 30 minutes left, the shots are divided into eight. V01 appearance 3 seconds, V02 turning the key 3 seconds, V03 walking to the shelf 4 seconds, V04 back 3 seconds, V05 returning to the shelf 4 seconds, V06 profile 4 seconds, V07 window and rain 3 seconds, V08 pulling inside the store 6 seconds.

Place only one main action on each card. For V05, "Put the book on the shelf with your right hand and let go." Set the camera to a fixed medium. If you add in things like walking here, looking up, and having the camera wrap around you, you won't know which mistakes you should fix. There are only three passing conditions for each card. Natural contact with the person's glasses and apron, the binding of the book, and the shelf. There is no need to use the same scoring sheet for cuts with different intentions.

Order of creating candidates

Start by creating multiple short, low-cost candidates for each cut, storing the seed or equivalent reproduction key, model version, input references, prompts, and output IDs. Write the description in the following order: subject, action, location, light, and camera. Rather than piling on “cinematic” and “beautiful,” increase the facts on screen. Use start/stop frames and camera controls only if you are sure the service supports them. Google's Veo official page provides information on references, start/end frames, extensions, and audio, but the actual region, plan, and API availability should be double-checked in the official document at the time of production.

  1. 13 suggestions for each shot
  2. 2Verify identity with still image
  3. 3Verify movement and contact
  1. 1Fail
  2. 2Select one failure type
  3. 3Change only one input
  1. 1Pass
  2. 2Make the last frame a reference for the next cut
Consider the sequence and each role.

Categorize and fix failures

If a person becomes a different person, check whether the reference is only from the front, whether the face is too small, or whether the motion is too large before adding more facial-description text. If an object passes through another object, divide the motion into "holding," "moving," and "placing," and make the moment of contact a separate cut. If you fail to change the background, try again with a wide reference, a fixed camera, and a shorter length, and if it's still unstable, use an angle of view that doesn't show the background in the editing. This is not a defeat. The decision was made to absorb the uncertainty of production through the design of the screen that the audience reads.

Scoring will be 0 to 2 points each for identity, physical contact, background, camera intent, and time connection. Even if the total score is high, if the identity of the main character is 0, it will be rejected. On the other hand, small background differences are acceptable as long as they are not visible in the next cut. In production, rather than perfection, we judge whether errors will benefit the story.

Edit and final audit

First, connect V01 to V08 without sound. Look at which way the person walks, the position of the shelves, the direction of the light, and whether belongings are kept across the cut. Next, place the rain, interior ambience, and book-placement sounds on separate tracks. When using voice generation, check the meaning of the lines, mouth shape, timing of sound effects, and terms of use. As Sora's system card explains, attribution and security measures are not post-processing outside of production. At publication, separately audit real people, trademarks, misleading scenes, and provenance metadata.

Finally, summarize on one page just what you changed and what changed. Even if you only watch the completion video once, the failure chart and reference pack will speed up the next 30 seconds. The value of short stories is not just the number of views. It's about being able to have a map of how the models and processes you handle break down.

Recording format for passing failures to the next attempt

Candidates are divided into "successful generation" and "adopted as video." The former is that the output was obtained, and the latter is that the requirements of story and continuity were met. Candidate number, version of input reference, model and version, length, ratio, main settings, generation time, scoring, and correction hypothesis are left. If Candidate A is beautiful, but the spine of the book has been changed, and Candidate B is plain, but the hands and book are natural, then choose B and decide whether you can edit only the color.

The sounds, the rain, the door, the paper, the room, the music, and the dialogue don't all come together as one. Determines how much reverberation is passed from the sound of the door in V03 to the sound of the paper in V04. If the entire sound cuts out every time the screen changes, the video will sound like a collection of generated clips. It is also effective to test by removing one cut from the final version. If the story passes, it's duplication, and if it doesn't, you can see whether the cut was responsible for line of sight, direction, emotion, or time.

  1. 1Candidate Output
  2. 2Technical Pass: Not Broken
  3. 3Story Pass: Will it lead to the next cut?
  1. 1Story Passed
  2. 2No Sound Edit
  3. 3Add Sound with Separate Track
  4. 4Public Audit
  1. 1Rejected
  2. 2Save Failure Label and Input Version
  3. 3Change one of the following conditions
Consider the sequence and each role.

Specific scoring chart for taking cuts

Using V05 "Putting the book back on the shelf" as an example, if the person's glasses and apron match the reference, score 2 points for identity, if only one is different, score 1 point, and if it is a different person, score 0 points. If a book touches your hand and naturally enters a gap in the shelf, it will give you 2 contact points, if there is something suspicious just along the way, you will get 1 point, and if it penetrates or proliferates, you will get 0 points. If the wood grain of the shelf or the position of the window are continuous, the background will be given 2 points, and if the camera is the specified fixed medium, the camera will be given 2 points. Rather than total points, place a gate that regenerates if identity or contact is 0. This is because the average score does not hide the failure that the audience notices first.

When making corrections, only one input is changed. If the contact is 0, separate walking, holding, and returning, or change the hand to a larger reference. If the background is 0, add a wide location reference and stop camera movement. If the camera is 0, reduce the movement and pass once at a fixed angle of view. If identity is 0, add profile and full-body references and shorten the motion duration. If you change the style word and resolution at the same time, you will misunderstand the reason for the fix.

Before exporting for publication, review the three speeds on the editing timeline. It's real time, visual without sound, and time to listen with just sound. Movements that can be understood without sound may become unclear when relying on sound. If the scene transition is unnatural with just sound, add a few frames of rain or room sounds. By separating cause and effect, which can be read from images alone, from emotion, which is supplemented by sound, the audience will not have to worry about the boundaries between cuts, even if the length is short.

Once you have finished making corrections, you can view not only the preview in the editing software, but also the actual exported file from beginning to end. It's hard to notice on the timeline that the sound is out of sync, the last frame is black, subtitles are cut off, and details are lost due to compression. The 30-second verification closes only after you confirm the final selection back to the output file.

Add a release gate, not another prompt pass

The 2024–2026 shift toward longer and more controllable clips makes a 30-second exercise more useful, not less: it exposes continuity and rights errors that a single demonstration hides. Add a final release gate with a cut-by-cut ledger: source reference, consent/licence status, generated file hash, edit decision, audio origin, visible watermark/provenance check, and owner. This is a production control, not a claim that any generator is safe by default.

The current month boundary yielded no primary release that changes this workshop's procedure. Keep the existing dated model notes as history and update this chapter only when a confirmed input, export, or provenance contract materially changes the gate.

Make the 30 seconds a deliverable with two masters

The September 23 HDR guide makes a practical addition to this workshop: the cut list must name its delivery target, not only its resolution. Keep an SDR review master as the approval artifact. When an HDR export is in scope and the selected tool supports it, make it a separately checked deliverable with its own display test; do not assume an HDR-looking preview is an HDR master.

Add two columns to the cut ledger: reference version and delivery transform. At the end of each cut, verify title/text legibility, audio boundary, reference consent, and color continuity. Render a low-cost editorial pass before a final pass, then compare the exact same frame at every cut. This turns 30 seconds into a verifiable sequence rather than a single generated file.

MENTAL MODEL / SHOT DESIGN

Give each shot one job.

01Establish the setting4 s
02Show the action4 s
03Leave a clear meaning4 s

Total: 12 seconds. Define a reference image, a start state, one action, and an end state for each shot. Arrange the still images in this order before generation to find transitions that do not communicate your idea.

SOURCES

01
Sora System Card ↗openai.com · 2024-12-09
02
Veo official model page ↗deepmind.google · unknown
03
OpenAI Sora system card ↗openai.com · 2024-12-09
04
aivideo discussion (community experience) ↗www.reddit.com · 2026-08-19
05
Runway product-video workflow ↗dev.runwayml.com · 2026-09-23
06
Runway HDR video guide ↗dev.runwayml.com · 2026-09-23
07
Runway fashion workroom workflow ↗dev.runwayml.com · 2026-09-24
08
Runway model guide ↗docs.dev.runwayml.com · unknown
09
Runway deprecated standalone tools ↗help.runwayml.com · 2026-08-05

YOUR NOTES