A believable frame is not an event record
At 2026-10-01 17:59:53 UTC (October 2 in Japan), the authors submitted the v1 preprint PROWBench. Its date-only source metadata is October 1. The paper header says ROWBench, while its abstract, project, and repository call the benchmark PROWBench; this note uses the latter name and does not infer a separate release. Its narrow question is whether a generated video visibly realizes an interaction that an executable program recorded. A gate may be recorded as opening and a person as passing through it, while a plausible-looking clip omits, delays, or changes that event.
The paper's design keeps two artifacts separate. An executable world record stores entity state, camera state, and timestamped events, including facts outside the camera view. A renderer then turns synchronized views of that record into inputs for a video model. The record gives the evaluator a reference for what should have happened; the generated clip supplies evidence for what was visibly rendered. That distinction is more useful than asking a single scorer whether the output looks good.
- 1Executable scene and event contract
- 2replayable states, camera path, event timeline
- 1Same record
- 2proxy or reference clip
- 3generated video
- 1Event window and expected end state
- 2evaluator decision: pass, reject, or unknown
- 1Visual plausibility score
- 2separate review; never a substitute for event evidence
The authors report benchmark results using synchronized proxy views and event/timeline checks beside geometric and perceptual measures. Those are author measurements, not an independent result from this library, a guarantee for a video service, or evidence that a generated video is physically correct.
Two questions a score should not collapse
The paper separates an event's required action/end state from whether an action appears in its assigned time segment. A clip can depict an action in roughly the expected window but miss its required end state. It can also place entities convincingly while violating the timeline.
This leads to a practical evaluation card for a single shot:
| Question | Evidence to retain | A result that must stay separate |
|---|---|---|
| Did the intended action appear in its time window? | event ID, start/end timestamps, expected action, selected frames | visual quality |
| Did the interaction reach the specified end state? | participants, end-state predicate, evidence frames | prompt compliance elsewhere in the clip |
| Is the output visually usable? | an explicit human review rubric | proof that the program's event occurred |
| Can the evaluator decide? | reason and missing evidence | a forced pass/fail guess |
unknown is a real evaluation outcome. Use it when a participant is occluded, the event is outside the clip, a reference timestamp cannot be mapped to the output, or a reviewer cannot distinguish the specified end state. Do not turn an ambiguous VLM response, a high similarity score, or a smooth motion curve into a pass. A human reviewer can also make mistakes, so retain the evidence frames and the event rule rather than only the label.
An unexecuted N=1 offline fixture
This original learning fixture is deliberately small and has not been run here. It requires no video-generation API, paid model call, hardware control, or upload of personal footage. Use a synthetic storyboard or a licensed internal animation with one event only: for example, actor A reaches the blue gate; the gate opens; actor A is beyond the gate by 5.0 s.
- Write one event row before generating or reviewing anything: event ID, participants, time window, expected action, end-state predicate, camera view, and an owner.
- Create a compact reference record:
0.0–1.0 approach,1.0–2.0 gate opens,2.0–5.0 actor beyond gate. Save the record and a proxy drawing locally. - If a clip already exists under an authorized local licence, select fixed frames around each boundary. Do not submit it to an external service for this exercise.
- Review the action and end state against the event row. Mark
passonly when both are visible under the written rule; markrejectwhen the rule is visibly violated; otherwise markunknownand name the missing observation. - Keep visual-quality notes in a separate column. A beautiful rejected clip is evidence of a rendering failure, while an ugly event-faithful clip may point to an art-direction issue instead.
The fixture stops if the record lacks a timestamp mapping, the source licence does not authorize the review, participant identity cannot be evaluated, or the end-state rule is vague. It does not repair a prompt, retry a model, call an agent, or promote a result to a real-world claim. Its output is an auditable row, not a benchmark leaderboard.
Limits that change the next experiment
The full 39-page paper describes limits that matter for this fixture: its event/appearance checks use a VLM judge not calibrated against human annotation; camera measures inherit pose-estimator error; and long-horizon reappearance plus independent multi-view identity remain unresolved. Preserve unknown, use a human review for high-stakes footage, and do not generalize one event row to a whole model.
A dated VBench++ v1, submitted on November 20, 2024, provides a useful preceding contrast: it organizes video quality and condition-consistency dimensions, whereas this fixture asks for evidence against one executable event record. They answer different questions; neither score replaces the other.
The paper's project and repository are useful implementation references, but their present pages do not grant access to a provider account, establish a production licence for other media, or replace a local security and retention review. Connect this fixture to reference-first production: references specify appearance and intention; an event contract specifies what a clip must visibly prove. Keep both before asking whether a shot is ready to edit.
MENTAL MODEL / SHOT DESIGN
Give each shot one job.
Total: 12 seconds. Define a reference image, a start state, one action, and an end state for each shot. Arrange the still images in this order before generation to find transitions that do not communicate your idea.
Sources
Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.
01