Video is more difficult than “adding movement to images”
A beautiful still frame does not imply that the next frame will also be coherent. Between two frames lie a person’s hands, folds of clothing, the light source, and the camera position. Video-generation failures stand out because temporal cause and effect breaks down: a hand that set down a cup gains fingers, or a window changes as the camera pans. Viewers immediately notice differences that might pass unnoticed in a still image because they read them as broken continuity.
Early image generation competed with GANs, autoregression, and diffusion. The scaffolding for understanding current large-scale video models is latent diffusion. First, an encoder projects a high-resolution video into a compressed latent space. Next, the model gradually returns the "latent blocks containing time" that have become noise to images that match the conditional statement and reference image. OpenAI's Sora system card compresses video into a low-dimensional latent representation and splits it into spatiotemporal patches. The important thing here is that the patch spans not only screens but also time. Even in scenes where a person disappears into the distance and returns, the model has clues to maintain the same character.
- 1Text Reference Camera Intent
- 2Conditional Expression
- 1Video
- 2latent compression in VAE
- 3Spatiotemporal patch
- 4DiT/Transformer
- 1noisy latent
- 2Repeated Restoration or Vector Field Tracking
- 3latent video
- 4Decoding
Diffusion, Flow and DiT are not three competing names
Diffusion is an idea of learning inverse operations that add noise little by little to data. When generating it, it starts from a random number and updates it several dozen times to get closer to the image. It was the first major stepping stone from images to video because it was stable and easy to connect to conditioning and editing. On the other hand, the number of iterations drives up speed and cost. For longer videos, the number of spatiotemporal tokens increases, and the computational complexity of the attention mechanism becomes heavier.
Flow Matching is a framework that directly regresses the "velocity vector from the current point toward the data distribution." Original paper states that continuous normalization flows can be scaled up by learning without simulation, and diffusion paths can be included as a special case. The reason the producers know this is not to chase new advertising words. This is because the number of steps, distillation, and real-time performance of the sampler determines how many samples can be sampled. Fast models with short iterations are good for exploration, but speed does not guarantee long-term causality or completeness of detail.
DiT is an abbreviation for Diffusion Transformer, and is a design that replaces the noise remover from a convolution-based network with a Transformer-based one. Images, audio, and videos can be treated in the same way as patch sequences, making it easy to scale and natural to combine with text conditions. Sora's idea of creating patches of visual data of variable length, resolution, and aspect ratio does not constrain material to one fixed frame size. However, patching is not magic that embeds the laws of physics. Plausible gravity, liquids, and contact are predictions of learned distributions, not results of mechanical simulations.
What has changed and what remains
The strength of video modeling is that it allows you to create storyboards, product images, and worldview verification before shooting. The official LTX-Video repository divides small distilled versions into quick iterations and larger versions into quality ones, and also presents control models for Pose, Depth, and Canny. LTX-Video's Apache-2.0 is a code license, so do not assume that the model-weight and associated-material terms are the same. It is necessary to read the terms and conditions of each model-weight distribution page separately.
For the same reason, "open" does not promise local execution. model-weight availability, disclosure of inference code, commercial use, training use, and handling of portraits are separate axes. In the cloud provision model, there is a difference between being able to use the API and making learning data and internal evaluations fully public. Benchmarks from model providers can serve as material for selecting candidates, but they do not constitute an independent evaluation that includes the subject, angle of view, sound, and editing process.
Minimal experiment for moving hands
Readers who have not yet run the model will start with a comparable exercise. Fixed only 5 seconds, 24fps, 16:9 "slowly panning from left to right of a steaming cup by the window". For the prompt, separate the subject, location, light, camera, movement, and change you want to prohibit into one sentence each. Record the seed, reference image, resolution, steps, CFG, generation date and time, and model version in the table. Next, change the camera from fixed to horizontal movement just once. Don't just say "I like it" to see if it's improved, but instead score 0-2 points for person identity, contact, background stability, fingers, camera intent, and sound synchronization.
- 1Identical reference and seed
- 25 seconds of fixed camera
- 3score
- 1Change only one condition
- 25 seconds of lateral movement
- 3record the difference
- 1Causes of low scores
- 2categorized by prompt or reference or control or edit
What we get from this experiment is not the "best video." Every model has conditions under which it will fail, and the types of failures it can tolerate in its process. If the generator is viewed as a replacement for photographic equipment, the failure appears to be accidental. If we treat it as a probabilistic candidate generator and divide cutting, selection, correction, and editing, changes in technology will translate into speed of production.
Notes on evaluation
Sora's public materials describe the video as a diffusion model and also touch on its provenance and security measures, including C2PA. However, the provision format, upper limit, and usable inputs at the time of publication may change due to product revisions. This text is a conceptual explanation based on materials obtained on 2026-10-04, and is not a definitive statement on the availability or cost of a specific model. Before handling human portraits, existing works, or news-like footage, check each service's policies, consent, and display/provenance requirements.
Connect generator internals and production decisions
Spatiotemporal patches make it easier to handle relationships over a wide time range, but they do not determine the importance of scripts or the accuracy of product names. The creator cuts long movements into short intentions and clearly indicates the entrance and exit of each intention. Generation settings are determined in the following order: scale and proportion, reference and composition, movement, style, and finally resolution. We use low-resolution rough sketches to determine who moves, where, what, and in which direction, and improve the surface quality with only passing cuts.
- 1Low cost composition test
- 2Identity and action test
- 3Connection test
- 1Pass Cut
- 2High Resolution and Compensation
- 3Edit Timeline
- 1Issues found in editing
- 2Revert to input criteria
- 3Update the following design
In the evaluation, we not only consider the best candidates, but also the rate of successful candidates. Count how many times out of 10 the identity, contact, camera, and background conditions are met. Even if you have a masterpiece, if the selection rate is low, you will not be able to predict the delivery date and cost. Models that are less flashy but have a high pass rate under certain conditions are strong in series. Comparison is not an aesthetic evaluation of a one-time competition, but a process of comparing the probabilities of processes.
This record can be re-evaluated with the same input set during future model updates. This will serve as a reference line to determine whether it is the model that has changed or the production conditions that have changed.
2024–2026 boundary: architecture claims are not product availability
The 2024 Sora research release made patchified video/image representations and diffusion transformers a useful mental model for production planning. But architecture continuity does not guarantee a continuing product contract: OpenAI’s Sora page now says that the product became unavailable on 2026-04-26. Treat a model paper, a system card, and a live product surface as three separate observations. For a new workflow, retain the date on which each was checked and plan an exportable intermediate format rather than depending on a named service.
An August 19, 2026 community post about stitching versus one long generation is a useful reminder that users still diagnose shot boundaries as a workflow problem. It is not evidence about a model's capability. Test this yourself by rendering one action as a single clip and as two cuts, then compare identity, object position, and editability frame by frame.
September 2026 production context: architecture is not the delivery contract
Runway's September 23 product-video guide presents a current workflow surface with multi-shot planning and a model-selection table. It is an official vendor workflow guide, not evidence that the listed models share architecture, provider terms, or measured quality. Add a delivery layer to the architecture map: first choose a bounded shot task, then record the source/reference, generated intermediate, edit decision, export codec and review result. A “world model” claim remains separate from whether a workflow can preserve identity and artifacts through editing.
Exercise: make two 4-second clips from the same reference sheet, one with a fixed camera and one with a planned cut. Before rendering, state which variables must remain invariant. After export, inspect the boundary frame, audio start, color transform, and provenance metadata. The failure record belongs beside the latent-model explanation.
MENTAL MODEL / SHOT DESIGN
Give each shot one job.
Total: 12 seconds. Define a reference image, a start state, one action, and an end state for each shot. Arrange the still images in this order before generation to find transitions that do not communicate your idea.
SOURCES
01YOUR NOTES