Before the comparison table, separate the provision formats
There are differences between video models that cannot be compared using the word "latest." Cloud products can offer large-scale computation and a unified UI, but are subject to price, latency, geography, terms of service, and updates. Open weights can enable experiments on-site or in your own cloud, while taking over the GPU, dependencies, inference time, model-weight licensing, and operational responsibility. A public repository may present the inference code but does not imply model-weight access or commercial-use rights.
- 1Requirements: Voice/Reference/Speed/Confidentiality/Budget
- 2Choose delivery format
- 1Cloud
- 2API, UI, Policies, Pricing
- 1open weights
- 2Check the model-weight license, VRAM, and reproduction procedures
- 1Common
- 2Evaluate your use with a fixed test set
Google's Veo official page explains character consistency from reference images, style references, camera control, start/end frames, scene extensions, and audio. Preference ratings posted by providers are useful as a hypothesis of ability, but it is necessary to read the comparison conditions, duration, resolution, and presence or absence of audio. For example, being selected for a short comparison does not guarantee the continuity of an edited 30-second piece.
OpenAI's Sora 2 system card was published on September 30, 2025; it explains video and audio, physicality, and controllability, and covers safety measures against use without the consent of a real person or use that may lead to misunderstanding. This is a security document, not a contract that says all users have access to the same inputs, outputs, and APIs. Review current product pages, regions, plans, and policies before publishing.
Wan and LTX: Local freedom comes with the responsibility of verification
Wan2.1 official repository shows the inference procedure and model group as an Alibaba-based open video generation project. Before creating an executable environment, read the README, model cards such as Hugging Face, license, required VRAM, sample resolution and length. GitHub star numbers and video demos do not prove the effective speed, stability, or commercial viability of a specific GPU.
LTX-Video provides distilled models for fast iterations, large quality-oriented models, ComfyUI, and an entry point to control and LoRA training. The Apache-2.0 repository is a code display. Do not assume the same weights, external components, LoRA used, or materials included in the output. Local inference can be an option that does not send input to an external API, but confidentiality can only be discussed after checking the data path, including logs, shared storage, and where the model is retrieved.
Turn selection into production testing
Using the same 8 shot list, experiment with people references, product references, start/end frames, camera movement, contact, audio, and aspect ratio. Quality is scored on character consistency, physics, camera, editing connection, text/logo inclusion, and audio synchronization. The action measures the time required per candidate, failure rate, retry cost, output storage, setting reproduction, and license confirmation time. Even if a model is excellent on its own, if it takes three hours to modify it, it may fail in daily production.
Initial implementations will start with non-confidential dummy material, short lengths, low resolution, and limited output. Leave human faces, existing IP, and news-like scenes for later, and check for necessary consents, provenance, labels, and terms before publishing. In rapidly evolving fields, a mechanism that allows the same set of evaluations to be re-executed at update intervals, such as 12 hours or monthly, is more effective over the long term than relying on fixed model names.
Translate official announcements into working instructions
References, start/end frames, camera controls, and scene extensions on Veo's official page each correspond to separate failures. Try character references for character identity, start/end frames for transitions between two compositions, camera control for screen movement, and scene expansion for length extensions. If you turn on all four functions at the same time and it doesn't work, you won't know which one is causing the failure. We'll start with a small test using just one feature to see what's different from standard text generation.
Sora's system card not only handles the ability to generate video and audio, but also handles safety aspects such as use without the consent of a real person, generation that leads to misleading identification, and restrictions regarding minors. The more physical the footage becomes, the more likely it is that viewers will mistake it for live action. Materials dealing with social events or people are checked for purpose, consent, labeling, and publication context before being considered for quality assessment. Safety measures are not just the job of model providers, but are part of the production process that involves selecting and publishing materials.
In Wan and LTX, the closer you go to local execution, the more important reproducibility records become. Store weight sources, commits, dependent libraries, GPU, quantization, workflow JSON, inputs and outputs together. If you only measure processing speed, you will not be able to produce the same output at a later date, and the comparative value will be lost. Only by recording the differences in your environment can you verify whether the lightweight version is really faster for your production.
Later availability check
This is a 2025 comparison note, not a current availability table. A later official check found that OpenAI says the Sora product became unavailable on 2026-04-26. Therefore do not select Sora from the 2025 entries as if it were a live option. Re-run the same matrix for each candidate's current account access, region, input/export contract, provenance features, and deletion behavior. Keep the old comparison as dated history.
No directly relevant official video-model release was confirmed in the 2026-09-04 to 2026-10-04 research boundary for the four named entries. Absence of a confirmed release is not a negative capability claim.
Replace the stale table with an access check, not a new ranking
The current Runway model guide lists an access surface on 2026-10-04, while its publication date is not supplied. Treat it as an undated current index: useful for discovering that Runway exposes multiple model choices, but not as a September release announcement or proof of access in another account or region. Runway's August 5 deprecation notice further says Gen-3 and Veo 3 standalone tools were deprecated on Runway; that is a Runway product-surface fact, not a statement that the original providers discontinued their models.
For every row in the 2025 matrix, add checked-at, provider, host surface, account/region observed, supported inputs, export format, and deprecation status. A third-party host listing MiniMax H3 Max, Wan 3.0, Seedance 2.5, Veo 3.1, or Gen-4.5 is availability information for that host, not a launch date, price, or capability measurement.
MENTAL MODEL / SHOT DESIGN
Give each shot one job.
Total: 12 seconds. Define a reference image, a start state, one action, and an end state for each shot. Arrange the still images in this order before generation to find transitions that do not communicate your idea.
SOURCES
01YOUR NOTES