GENERATE TOGETHER. COMMIT ONE.
Commit–One.
Decoupling Generation and Commitment
for Interactive Video Generation
Bidirectional model · 5-second generation chunks.
Predict a large chunk. Commit only what is needed.
Keep the rest of the future open to change.
Fast generation.
Still waiting to interact?
A new condition should change what the viewer sees promptly. Even a fast generator can leave it waiting for generation and playback boundaries.
Must fine-grained interaction require
fine-grained generation?
Smaller chunks offer more frequent updates, but can limit joint temporal modeling and parallel efficiency.
Generate together.
Commit one.
A prediction does not have to be a promise.
Jointly generate K latents, then deliver at q = 1. Protect the necessary prefix and keep future predictions replaceable.
Commitment ≠ playback. Committed output can no longer be replaced, even if it has not yet been displayed.
A reusable future.
A two-state controller.
Produce generates a chunk; Consume reuses predictions. New conditions trigger feasible suffix replanning; refill keeps the stream supplied.
Chunk size is a design choice.
Larger chunks generally improve Wan quality within the same resolution and step budget. Suitable parallelism can also keep them efficient—there is no single best K for every setting.
VBench total, K = 1 → 3.
Gains are not monotonic across settings.
K = 1 on one H200 vs. K = 7 on eight H200s, per forward.
One H200 → eight H200s, per forward. More GPUs do not always help.
Tripling Wan’s current-chunk latents (K = 1 → 3) raises single-GPU forward time by 2.05× at 480p and 2.74× at 720p. These profiles exclude text encoding, VAE decoding, and the sampler loop; they are not end-to-end response times.
Wan quality · all resolutions and step budgets
Separately trained K = 1/3/7/21 configurations; Toy6K training, 946 VBench prompts × five seeds, 81 frames at 16 fps. Total is the VBench aggregate; the remaining columns retain the reported component and auxiliary scores.
| Resolution | Steps | K | Total | Quality | Semantic | Dynamic | VisionReward | Instruction |
|---|---|---|---|---|---|---|---|---|
| 480p | 1 | 1 | 79.11 | 81.80 | 68.34 | 4.00 | 6.50 | 90.00 |
| 480p | 1 | 3 | 80.38 | 83.40 | 68.29 | 58.00 | 4.81 | 94.00 |
| 480p | 1 | 7 | 81.35 | 83.97 | 70.86 | 48.00 | 5.79 | 100.00 |
| 480p | 1 | 21 | 81.68 | 84.83 | 69.08 | 85.00 | 9.93 | 94.00 |
| 480p | 2 | 1 | 80.07 | 82.76 | 69.34 | 14.00 | 9.31 | 98.00 |
| 480p | 2 | 3 | 81.62 | 84.63 | 69.57 | 57.00 | 9.82 | 100.00 |
| 480p | 2 | 7 | 80.84 | 84.06 | 67.94 | 93.00 | 9.29 | 100.00 |
| 480p | 2 | 21 | 81.84 | 84.62 | 70.74 | 71.00 | 10.17 | 100.00 |
| 480p | 4 | 1 | 79.00 | 81.38 | 69.50 | 13.00 | 7.65 | 96.00 |
| 480p | 4 | 3 | 81.66 | 84.70 | 69.51 | 80.00 | 9.64 | 100.00 |
| 480p | 4 | 7 | 81.70 | 84.73 | 69.56 | 78.00 | 10.19 | 98.00 |
| 480p | 4 | 21 | 82.13 | 85.03 | 70.54 | 85.00 | 9.66 | 98.00 |
| 720p | 1 | 1 | 78.94 | 81.37 | 69.19 | 4.00 | 5.82 | 92.00 |
| 720p | 1 | 3 | 79.99 | 82.70 | 69.18 | 47.00 | 3.93 | 98.00 |
| 720p | 1 | 7 | 80.78 | 83.05 | 71.67 | 41.00 | 5.23 | 96.00 |
| 720p | 1 | 21 | 80.49 | 83.10 | 70.04 | 67.00 | 9.56 | 100.00 |
| 720p | 2 | 1 | 79.86 | 82.40 | 69.70 | 9.00 | 9.43 | 98.00 |
| 720p | 2 | 3 | 81.52 | 84.11 | 71.18 | 48.00 | 10.10 | 100.00 |
| 720p | 2 | 7 | 80.97 | 83.79 | 69.67 | 84.00 | 9.25 | 94.00 |
| 720p | 2 | 21 | 80.26 | 82.61 | 70.84 | 39.00 | 9.91 | 98.00 |
| 720p | 4 | 1 | 78.84 | 81.19 | 69.44 | 22.00 | 7.18 | 92.00 |
| 720p | 4 | 3 | 81.59 | 84.36 | 70.49 | 74.00 | 9.59 | 100.00 |
| 720p | 4 | 7 | 81.86 | 84.62 | 70.83 | 73.00 | 9.77 | 96.00 |
| 720p | 4 | 21 | 81.65 | 84.43 | 70.52 | 66.00 | 9.36 | 96.00 |
Wan efficiency · H200 forward profiles
Batch one, BF16, a shared reference checkpoint, and synthetic inputs with a 21-latent prefix. Slowest-rank median over 30 synchronized forwards after 10 warm-ups, including communication and output gathering. USP1/2/4/8 use 1/2/4/8 GPUs; all values are milliseconds.
| Resolution | K | USP1 | USP2 | USP4 | USP8 |
|---|---|---|---|---|---|
| 480p | 1 | 77.5 | 63.7 | 65.9 | 72.1 |
| 480p | 3 | 159.2 | 91.5 | 67.0 | 72.9 |
| 480p | 7 | 358.3 | 194.7 | 105.2 | 78.6 |
| 480p | 21 | 1406.2 | 750.0 | 384.1 | 211.0 |
| 720p | 1 | 218.7 | 143.4 | 95.8 | 72.7 |
| 720p | 3 | 599.6 | 318.4 | 166.5 | 119.4 |
| 720p | 7 | 1521.4 | 809.2 | 443.8 | 276.4 |
| 720p | 21 | 6432.8 | 3335.5 | 1737.1 | 899.7 |
LTX-2.3 shows why the choice is model-specific. Its six visual scores move differently with K, and small chunks do not benefit from more GPUs in these cached-forward profiles.
LTX-2.3 · visual quality and cached-forward profiles
Internal32: 32 matched cases/seeds, 361 frames at 832 × 448 and 24 fps; Internal8K training, one training run per K. Three DiT forwards with schedule [999, 757, 522, 0] (0 is the terminal boundary). Visual scores are ×100 and have no combined total. SC/BC: subject/background consistency; Aesth.: aesthetic quality; Imaging: imaging quality; Smooth.: motion smoothness.
| K | SC | BC | Aesth. | Imaging | Smooth. | Dynamic |
|---|---|---|---|---|---|---|
| 1 | 87.97 | 94.68 | 52.82 | 73.09 | 99.62 | 31.25 |
| 3 | 88.81 | 94.31 | 52.21 | 65.98 | 99.53 | 46.88 |
| 9 | 89.42 | 93.27 | 50.72 | 68.96 | 99.56 | 34.38 |
| 15 | 84.96 | 92.21 | 51.01 | 65.98 | 99.50 | 78.12 |
| 45 | 88.72 | 93.68 | 52.40 | 68.75 | 99.64 | 40.62 |
| K | USP1 | USP2 | USP4 | USP8 |
|---|---|---|---|---|
| 1 | 274.0 | 440.5 | 479.6 | 526.1 |
| 3 | 255.6 | 423.5 | 457.0 | 511.2 |
| 9 | 426.7 | 417.8 | 456.6 | 502.0 |
| 15 | 660.3 | 420.4 | 461.0 | 509.4 |
| 45 | 2016.2 | 1158.4 | 634.0 | 507.9 |
Keep chunk size available as a quality–efficiency choice, while making commitment fine-grained.
Make joint generation faster.
Keep interaction fine-grained.
Decoupling generation and commitment opens another route to responsiveness: accelerate high-quality joint generation.
It also motivates exploring bidirectional teacher models for interactive streaming.