How Far Are We from Probabilistically Aligned World Modeling?
Benchmark and protocol: PAWBench formalizes probabilistic alignment with 50 scenarios across eight physical mechanism groups. PAWEval maps repeated rollouts to terminal outcomes, builds empirical distributions from readable in-schema outcomes.
Model diagnosis: across eleven video generators, no model combines accurate outcome probabilities, broad support recovery, and reliable performance across scenes.
Distribution shaping: language requests named futures, initial noise sampling broadens finite-budget exploration, and fine-tuning shifts learned outcome frequencies. These interventions separate inference-time steering and exploration from changes to the model's learned predictive distribution.
1 Shanghai Jiao Tong University2 Shanghai AI Laboratory3 Krea AI4 Hugging Face5 Shanghai Innovation Institute6 Tongyi Lab7 The University of Hong Kong
† Corresponding authors.
Click to jump to each section.
How far are current video generators from probabilistically aligned world modeling?
Current video generators can follow instructions, maintain temporal coherence, and produce plausible continuations from an observation and an action-like prompt. These capabilities have strengthened the case for treating them as visual world models. Using them as world models requires testing the distribution induced by repeated rollouts.
A world model must do more than render visually plausible rollouts. When a process is stochastic, the same initial observation and action can lead to multiple physically valid futures. We call a model probabilistically aligned when the distribution induced by its repeated rollouts captures both which futures are possible and how likely each one is. Interaction and planning depend on both the possible consequences of an action and their relative likelihoods.
Existing evaluations largely treat generated videos as isolated outputs, scoring visual quality, temporal coherence, prompt alignment, or physical plausibility one video at a time. These criteria cannot reveal whether repeated generations collapse to a subset of valid outcomes or assign those outcomes the wrong probabilities. One plausible future is therefore not enough; PAWBench asks how far current generators are from probabilistically aligned world modeling.
Figure 1: One plausible future is not enough. PAWBench asks whether repeated rollouts recover the support and probabilities of possible futures.
What probabilistic alignment requires
Support alignment
Repeated rollouts should recover the distinct valid terminal outcomes available under the same condition. Producing only one or two outcomes across repeated samples indicates that the generator has missed part of the valid support for that specific scene.
Probability-mass alignment
When reference probabilities are known, each valid outcome should also occur in the right proportion. A model can cover every outcome and still allocate the wrong probability mass among those outcomes.
PAWBench evaluates the induced distribution over terminal outcomes. This abstraction groups together trajectories that reach the same result; it does not claim to measure the complete distribution over intermediate states or frames.
PAWBench: A Benchmark for Possible Futures
Each PAWBench item specifies a source image, one atomic action prompt, and a finite set of valid terminal outcomes. When available, it also provides a reference distribution over those outcomes. A generator receives the same image and action for K repeated rollouts, which PAWEval maps to terminal outcomes and aggregates into an empirical distribution. Fixing the inputs makes variation across rollouts reflect the model-induced future distribution rather than changes in the prompt or image. This repeated-rollout contract is shared by every system in the comparison.
The benchmark contains 50 scenarios across eight physical mechanism groups. PAW-Calibration includes 25 scenarios with analytically specified or symmetry-derived reference distributions and tests probability-mass alignment. PAW-Coverage includes 25 scenarios whose valid outcomes can be enumerated but whose relative probabilities cannot be specified reliably, and tests valid-support recovery. We report the tracks separately because recovering every valid outcome does not establish that the model assigns the right probability to each one. The two tracks answer different questions about probabilistic alignment.
Figure 2: PAWBench scenario taxonomy. PAWBench covers eight mechanism groups under fixed initial observations and actions. PAW-Calibration contains calibrated probability scenarios with analytically specified reference distributions, while PAW-Coverage contains complex stochastic interactions evaluated by valid-support coverage. Each scenario fixes the source image and action prompt, then scores repeated model rollouts by their terminal outcomes.
Repeated-rollout statistics are meaningful only when the future distribution is well defined. We therefore require a visible stochastic mechanism rather than prompt ambiguity, one atomic action whose completion can be judged from video, and a finite set of visually distinguishable terminal outcomes. Before evaluating any model, we finalize the source image, action prompt, outcome set, failure criteria, and, for PAW-Calibration, the reference distribution. We fix this scenario contract before any generator produces a rollout and before any result enters the benchmark.
PAWEval: Outcome-Level Evaluation for Repeated Rollouts
PAWEval is the outcome-level protocol that turns repeated video rollouts into a distributional measurement. Each scenario provides a fixed rubric describing the terminal physical result to inspect, the valid outcome labels, and the conditions for an outcome-readout failure. Gemini 3.5 Flash applies this rubric to each rollout and returns either an in-schema terminal outcome or a readout failure. The same rubric and outcome vocabulary apply across all K=50 rollouts for that scene.
PAWEval normalizes readable, in-schema outcomes into a conditional empirical distribution and reports readout failures separately because they are not physical outcomes in the scenario's valid set. PAW-Calibration compares this conditional distribution with its reference using total variation distance (TVD); PAW-Coverage measures the fraction of valid outcomes recovered. A model-scene pair is scoreable only when at least 20 of its 50 rollouts are readable and in schema. Scene Pass Rate (SPR) reports the percentage of all 25 scenes in a track that pass this gate. The conditional metrics and SPR therefore answer different questions and must be reported together for every model.
Figure 3: PAWEval converts repeated rollouts into empirical distributions over valid terminal outcomes, while reporting readout failures separately.
Evaluation on PAWBench
Evaluation Setup
We evaluate eleven current video generation models on both PAW-Calibration and PAW-Coverage. For each scenario, every system receives the same source image and action prompt and produces K=50 independent rollouts using its default inference configuration unless noted otherwise. A scene passes the outcome-readout gate when at least 20 of 50 rollouts yield readable, in-schema outcomes. Scene Pass Rate (SPR) is the percentage of all 25 scenes in a track that pass this gate. This keeps model comparisons on one common sampling and readout contract for all systems.
Benchmark Results
Table 1: Main PAWBench results. PAW-Calibration reports conditional TVD ×100 (lower is better); a 70/30 model distribution against a 50/50 reference scores 20, while an exact match scores 0. PAW-Coverage reports conditional valid-support recovery percentage (higher is better). Avg. is computed over passing scenes; SPR is the percentage of all 25 scenes that pass. N/A means that no scene in the reported mechanism group passes the gate.
Swipe horizontally to compare all metrics.
Main PAWBench results for eleven video generation models.
Model
PAW-Calibration: conditional TVD ×100 ↓
PAW-Coverage: conditional Coverage % ↑
Avg.
SPR
Toss
Rot.
Rout.
Draw
Avg.
SPR
Coll.
Stab.
Agent
Mat.
HappyHorse
43.1
92.0%
40.1
27.5
41.9
59.5
47.1
100.0%
45.0
36.8
49.0
58.3
Veo3.1 Fast
35.4
88.0%
37.1
12.2
32.4
55.0
41.8
100.0%
39.8
21.1
53.3
44.2
Kling 3 Std.
34.9
92.0%
29.9
20.3
38.7
46.0
52.8
88.0%
52.7
40.0
44.3
77.5
Seedance 2
30.5
100.0%
32.7
15.1
31.1
40.5
50.9
84.0%
48.6
50.0
44.3
67.5
Wan2.7
26.3
92.0%
26.0
12.2
26.8
36.5
50.0
96.0%
60.3
50.5
37.9
53.3
Wan2.2
26.3
64.0%
22.4
15.8
23.2
45.9
63.4
92.0%
76.8
56.7
55.0
58.3
LTX-2.3
30.1
24.0%
29.3
11.0
N/A
41.0
71.7
72.0%
82.7
50.0
59.5
90.0
LTX-2.5
30.2
60.0%
30.2
17.1
39.9
27.1
57.4
88.0%
66.7
30.0
59.4
57.5
Cosmos 3 Super I2V
20.5
80.0%
19.4
3.6
17.7
35.5
55.2
92.0%
56.7
31.7
49.4
81.7
LingBot-Video-MoE
41.8
72.0%
31.8
13.4
44.4
60.8
58.8
100.0%
62.2
61.8
46.5
72.5
MiniMax H3
24.2
68.0%
22.0
14.0
28.9
26.5
48.7
92.0%
59.5
39.5
36.4
58.3
Current video generators remain far from probabilistically aligned world modeling. Cosmos 3 Super I2V has the lowest conditional Calibration TVD, 20.5, but only 80.0% of its Calibration scenes pass the outcome-readout gate. LTX-2.3 has the highest conditional Coverage average, 71.7, but that average is computed over the 72.0% of Coverage scenes that pass. Seedance 2 and LingBot-Video-MoE pass every scene on Calibration and Coverage, respectively, yet neither leads the corresponding conditional metric. Conditional alignment and scene-level reliability must therefore be read together when comparing the eleven generators in Table 1.
Recovering valid futures is not the same as assigning them the right probability. A generator may expose many plausible outcomes while allocating probability mass incorrectly, or closely match the frequencies of observed outcomes while leaving valid alternatives unseen. PAW-Coverage and PAW-Calibration therefore expose complementary failures and remain separate rather than being collapsed into one score. Recovering valid support is necessary, but it cannot establish correct probability allocation within that support under the same condition in PAWBench.
Causal and Non-Causal Controls
A probabilistic world model should change its future distribution when the physical transition changes and preserve it when the transition does not. We test this with paired interventions: causal interventions alter the transition and its reference distribution, while non-causal interventions change only an irrelevant visual or textual cue that leaves the physical process unchanged. The paired design tests both requirements under the same source action.
Figure 4: Models underreact to physically causal interventions and overreact to non-causal cues. Upper and lower bars show outcome distributions before and after intervention; the pencil tilt changes the physical transition, while the Galton-board text leaves it unchanged. Bars are conditional on valid readouts, and this figure is diagnostic rather than a pooled benchmark score.
Model distributions shift incompletely or in the wrong direction under physically causal interventions. Under non-causal interventions, distractor text redirects probability mass even though the reference distribution is unchanged. The models respond to altered inputs, but their response does not reliably track whether the underlying physical transition changed or remained unchanged. This asymmetry persists under the fixed source action and readout rubric.
Robustness of the Evaluation Protocol
The main results could reflect either model behavior or a limitation of the evaluation protocol. We therefore test two measurement explanations for the observed gap: too few rollouts and disagreement between PAWEval and human judgments on clear terminal outcomes. Together, these probes address the sampling and readout stages while preserving the original benchmark target and outcome space used for the main model comparison.
Rollout Budget Sensitivity
Figure 5: Larger rollout budgets increase coverage but leave calibration largely unchanged. Conditional TVD and valid-support coverage are shown as the rollout budget grows from K=1 to K=100. K=50 is a shared evaluation budget, not a claimed convergence point.
Doubling the rollout budget leaves calibration largely unchanged, although coverage continues to rise for three of four models. Additional samples can uncover more valid outcomes without correcting their observed relative frequencies. Across eleven generators, observed TVD averages 31.2; in a matched finite-sample null that preserves each model's passing scenes and readable counts, 99% of simulated averages remain below 9.22, far below the calibration error observed across the evaluated generators under the same outcome-readout protocol.
Agreement with Human Judgments
We collect seven independent human judgments per video using the same scene-specific outcome space as PAWEval. The comparison below includes only videos for which both PAWEval and the human panel provide a clear terminal-outcome label. This gives both evaluators the same finite outcome vocabulary for every comparison.
PAWEval-human agreement on clear comparable videos.
Human-consensus subset
Videos
Exact matches
Agreement
Overall
888
722
81.3%
Low consensus (3-4 votes)
213
125
58.7%
Moderate consensus (5 votes)
154
119
77.3%
High consensus (6-7 votes)
521
478
91.7%
PAWEval matches the decisive human label on 722 of 888 clear, comparable videos (81.3%). Disagreement on clear outcomes therefore cannot by itself account for the PAWBench gap. The denominator excludes videos without a clear label from either side, so this comparison does not measure all-row accuracy or validate TVD, Coverage, or model rankings. It only tests agreement where both evaluators return a decisive terminal label.
Toward Probabilistically Aligned World Modeling
Generating one plausible future, or controlling which single future is produced, does not amount to controlling a distribution of futures. We intervene at three interfaces that can shape that distribution: language specifies a requested future, initial noise sampling selects among accessible futures, and model learning changes how probability mass is allocated among outcomes. The three probes separate inference-time steering and finite-budget exploration from changes to the learned predictive distribution. Each probe tests a different source of change in the future distribution.
Predicting and Steering Outcome Distributions via Language
Language can steer a video generator by naming a future in its prompt, but aligning repeated rollouts requires both selecting targets with the right distribution and realizing each selected target in video. We evaluate these requirements separately so that an error in future selection is not confused with an error in video realization. A failure at either stage leaves the repeated-rollout distribution misaligned even if the other stage succeeds.
We test three settings. Direct VLM future sampling repeatedly predicts what may happen from the same observation and action, then maps those predictions into the PAWBench outcome space without video generation. Prompt engineering (PE) asks GPT-5.5 to select a possible outcome and write a generator prompt for each rollout. Oracle PE supplies the target outcomes directly, following the reference distribution in PAW-Calibration and balancing valid outcomes in PAW-Coverage, to isolate whether the generator realizes the requested future. This ordering separates target selection from target realization under the same outcome contract for each repeated rollout shown here.
Swipe horizontally to compare all metrics.
Table 2: VLM distributions over possible futures. Repeated responses are mapped to PAWBench outcomes without video generation.
Model
PAW-Calibration: TVD ×100 ↓
PAW-Coverage: Coverage % ↑
Avg.
SPR
Toss
Rot.
Rout.
Draw
Avg.
SPR
Coll.
Stab.
Agent
Mat.
Qwen3.5 Plus
40.9
84.0%
36.4
15.1
44.9
59.0
35.9
96.0%
19.8
49.5
48.6
29.2
GPT-5.5
42.3
100.0%
35.6
42.9
41.9
56.0
34.3
100.0%
24.9
29.6
44.7
39.2
GLM-5V Turbo
34.8
88.0%
24.1
36.9
33.9
50.8
39.9
100.0%
30.9
56.6
42.3
38.3
Kimi K2.6
38.2
96.0%
32.6
32.1
36.8
58.5
45.1
100.0%
44.1
55.7
46.5
34.2
Gemini 3.5 Flash
39.4
100.0%
37.0
29.0
40.6
52.0
46.6
96.0%
49.5
39.5
48.9
43.3
Swipe horizontally to compare all metrics.
Table 3: Predicted outcomes do not substitute for target outcomes. The first row scores GPT-5.5 PE selections before synthesis; matched Base, PE, and Oracle PE conditions cover 25 scenes per track at K=50.
Model / condition
PAW-Calibration: TVD ×100 ↓
PAW-Coverage: Coverage % ↑
Avg.
SPR
Toss
Rot.
Rout.
Draw
Avg.
SPR
Coll.
Stab.
Agent
Mat.
GPT-5.5 PE
44.3
100.0%
40.1
49.5
43.6
49.0
35.0
100.0%
35.4
34.6
35.5
33.3
Wan2.2
26.3
64.0%
22.4
15.8
23.2
45.9
63.4
92.0%
76.8
56.7
55.0
58.3
+PE
27.1
92.0%
23.5
15.4
28.8
39.5
61.9
100.0%
68.1
50.7
62.9
57.5
+Oracle PE
15.3
84.0%
10.5
13.9
14.0
27.3
76.2
96.0%
71.1
72.7
82.8
76.7
Cosmos 3 Super I2V
20.5
80.0%
19.4
3.6
17.7
35.5
55.2
92.0%
56.7
31.7
49.4
81.7
+PE
31.6
92.0%
27.0
15.0
34.9
47.0
64.9
100.0%
71.1
62.0
53.5
76.7
+Oracle PE
12.8
88.0%
17.6
9.2
13.2
5.1
87.3
100.0%
84.8
91.4
91.3
80.8
MiniMax H3
24.2
68.0%
22.0
14.0
28.9
26.5
48.7
92.0%
59.5
39.5
36.4
58.3
+PE
38.9
96.0%
32.5
35.0
42.2
48.6
49.8
96.0%
59.3
41.8
40.5
57.5
+Oracle PE
10.8
96.0%
21.6
6.0
2.1
12.6
82.9
96.0%
82.4
91.4
82.1
76.7
LTX-2.5
30.2
60.0%
30.2
17.1
39.9
27.1
57.4
88.0%
66.7
30.0
59.4
57.5
+PE
36.4
96.0%
34.3
28.7
36.6
48.0
55.6
100.0%
59.8
42.1
51.4
68.3
+Oracle PE
18.0
88.0%
26.2
21.8
15.0
5.7
73.9
96.0%
56.0
78.0
89.9
73.3
Figure 6: Oracle PE often misses requested outcomes. Across the complete 2,500-row target-declaring roster per model, invalid, missing, unreadable, malformed, and different outcomes count as non-following.
Direct VLM future sampling is already misaligned: the best Calibration TVD is 34.8, and the best Coverage is 46.6%. GPT-5.5's PE selections score 44.3 Calibration TVD and 35.0% Coverage before video generation. Oracle PE improves conditional TVD and Coverage for every generator, but the generators realize only 37.6–58.1% of requested outcomes. Language can change individual rollouts without reliably aligning their aggregate distribution. Better target selection alone therefore does not align the generator under the fixed observation and action.
Broadening Outcome Exploration via Initial Noise
The language interventions change which future is requested. To test what sampling alone can recover, we instead hold the request and generator fixed and ask whether independent noise draws repeatedly visit the same modes even when other valid outcomes remain accessible. This isolates finite-budget exploration under one unchanged generator and prompt. The generator, prompt, and rollout budget remain unchanged throughout this comparison.
Couple to Control (C2C) introduces negative dependence among the K=50 initial-noise samples while preserving each sample's standard Gaussian marginal. C2C and independent sampling use the same action prompt, generator, rollout budget, and evaluation contract. This matched comparison isolates finite-budget exploration; it does not test a change in the model's learned distribution. Only the dependence among the sampled noise vectors changes.
Swipe horizontally to compare all metrics.
Table 4: Coupled noise broadens finite-budget exploration. Base and C2C use matched prompts, generators, rollout budgets, and readout contracts.
Model / condition
PAW-Calibration: TVD ×100 ↓
PAW-Coverage: Coverage % ↑
Avg.
SPR
Toss
Rot.
Rout.
Draw
Avg.
SPR
Coll.
Stab.
Agent
Mat.
Wan2.2
26.3
64.0%
22.4
15.8
23.2
45.9
63.4
92.0%
76.8
56.7
55.0
58.3
+C2C
25.7
84.0%
17.4
22.3
21.6
48.7
69.2
88.0%
82.4
57.5
59.4
68.3
LTX-2.3
30.1
24.0%
29.3
11.0
N/A
41.0
71.7
72.0%
82.7
50.0
59.5
90.0
+C2C
19.9
32.0%
18.6
7.2
34.0
27.7
74.8
72.0%
92.8
87.5
57.1
76.7
Cosmos 3 Super I2V
20.5
80.0%
19.4
3.6
17.7
35.5
55.2
92.0%
56.7
31.7
49.4
81.7
+C2C
19.4
72.0%
25.8
6.1
14.3
24.2
63.9
92.0%
74.9
56.7
49.1
76.7
Across the scenes that pass in each condition, C2C lowers mean Calibration TVD and raises mean Coverage for all three generators. The gains vary by mechanism and do not consistently raise SPR. Under this matched setup, C2C helps 50 rollouts explore the model's accessible outcomes more broadly; it does not establish that the model has learned a different probability allocation. The comparison remains finite-budget evidence, not a claim about the model's limiting distribution.
Reshaping the Model's Intrinsic Distribution via Fine-Tuning
Unlike language and noise interventions, changing the training distribution can alter a generator's learned allocation of probability mass. We train five Wan2.2 LoRA models on mixtures containing different proportions of left- and right-falling pencil videos while holding the training recipe and budget fixed. Each model is evaluated on upright and left-leaning pencil scenes with a direction-neutral prompt and K=50 rollouts. The two evaluation scenes therefore demand different state-conditional behavior from the same adapted model under a fixed action.
Swipe horizontally to compare all columns.
Table 5: Training mixtures reshape outcome mass. TVD ×100 for the Base model and five LoRA models trained on increasing proportions of left-fall examples.
Scene
Reference L/R
Base
LoRA training P(left) %
0
20
50
80
100
Upright
50/50
17.3
50.0
17.6
23.5
48.0
50.0
Left-leaning
100/0
41.3
97.2
50.0
18.9
0.0
0.0
Figure 7: Training steers outcome mass across scenes. Frequencies are conditional on readable left/right outcomes; invalid readouts are excluded.
Increasing the share of left-falling videos in training raises the generated left-fall frequency in both scenes, but the response is nonlinear and does not track the training proportion one-to-one. The result shows that training composition can reshape a model's outcome distribution, although the actual training distributions of current generators are unknown and are not inferred here. This experiment does not recover or estimate the pretraining distribution.
None of the five adapted models matches both scene-specific references. Mixtures that bring the left-leaning pencil to its 100/0 reference make the upright pencil fall left almost every time, instead of matching its 50/50 reference. The same global adjustment moves both scenes in the same direction. Probabilistic alignment requires the outcome distribution to change with the initial physical state under a fixed action. The adapted model instead applies a similar shift to both states.
Model Rollout Gallery
The gallery compares eight video generators under the same source image and action for eight representative PAWBench scenes: four PAW-Calibration mechanisms and four PAW-Coverage mechanisms. Each tab keeps the model order fixed and shows one base-condition rollout per model. Each displayed clip illustrates one sampled future rather than a model-level score or distributional result. The gallery supports qualitative side-by-side inspection of individual clips. Model ranking and distributional claims require the full repeated-rollout results reported above.
Calibration
Coverage
A-01 · stochastic toss
Coin flip
Action:Flick the coin once.
Base condition · rollout r000 · qualitative comparison
HappyHorse
Veo3.1 Fast
Kling 3 Std.
Wan2.2
Cosmos 3 Super I2V
LTX-2.5
LingBot-Video-MoE
MiniMax H3
Conclusion
World modeling requires more than producing plausible continuations: a model should capture the distribution of possible futures under the same initial observation and action. PAWBench makes this requirement measurable through repeated video rollouts. Across eleven current generators, no model consistently matches the reference probabilities while recovering the range of valid futures, and the gap cannot be explained by finite sampling or disagreement with human judgments on clear outcomes. Controlled interventions further show that model distributions do not reliably track causal changes in the physical process under the fixed PAWBench contract and readout rubric.
Language can request individual futures, coupled noise can broaden finite-budget exploration, and fine-tuning can shift outcome frequencies, but none reliably recovers the scene-conditioned distribution over possible futures. Plausible, diverse, or controllable rollouts do not by themselves establish probabilistically aligned world modeling across different physical states. The missing capability is state-conditioned control of the full future distribution.
Acknowledgment
We thank Rapidata for the human-study platform and the Cambrian authors for the webpage template.