INTERACTIVE VIDEO GENERATION

GENERATE TOGETHER. COMMIT ONE.

Commit–One.

Decoupling Generation and Commitment
for Interactive Video Generation

INTERACTIVE DEMO

Bidirectional model · 5-second generation chunks.

Predict a large chunk. Commit only what is needed.
Keep the rest of the future open to change.

Large-chunk generation · Fine-grained commitmentExplore the idea

Fast generation.
Still waiting to interact?

A new condition should change what the viewer sees promptly. Even a fast generator can leave it waiting for generation and playback boundaries.

PAPER FIG. 02 / Where response latency comes from
R waits to generate; C includes DiT sampling and required VAE decoding; Q waits to display. A schematic schedule, not measured timing.
THE QUESTION

Must fine-grained interaction require
fine-grained generation?

Smaller chunks offer more frequent updates, but can limit joint temporal modeling and parallel efficiency.

Generate together.
Commit one.

A prediction does not have to be a promise.

Jointly generate K latents, then deliver at q = 1. Protect the necessary prefix and keep future predictions replaceable.

PAPER FIG. 01 / Generation and commitment
Commit-One keeps large-chunk generation and replaces only the unprotected future. Positions represent video content, not measured latency.

Commitment ≠ playback. Committed output can no longer be replaced, even if it has not yet been displayed.

A reusable future.
A two-state controller.

Produce generates a chunk; Consume reuses predictions. New conditions trigger feasible suffix replanning; refill keeps the stream supplied.

PAPER FIG. 03 / A case through the controller
Camera pans left → right, K = 4, q = 1. Follow the state above and the output below: reuse, replace a feasible suffix, then refill. Playback continues in both states. Illustrative case; timing is schematic.

Chunk size is a design choice.

Larger chunks generally improve Wan quality within the same resolution and step budget. Suitable parallelism can also keep them efficient—there is no single best K for every setting.

CHUNK-SIZE ANALYSIS / WAN1.3BEnlarge ↗
Left: Wan four-step VBench total across K=1,3,7,21 at 480p and 720p. Right: 480p H200 single-forward latency for USP1, USP2, USP4 and USP8. Larger chunks benefit from suitable parallelism.
Left: four-step VBench quality at two resolutions. Right: 480p single-forward time on 1/2/4/8 H200 GPUs (USP). Lines connect tested settings; they do not define a quality–latency frontier.
QUALITY · WAN 480p / 4 STEPS79.00 → 81.66

VBench total, K = 1 → 3.
Gains are not monotonic across settings.

PARALLELISM · WAN 480p77.5 ≈ 78.6 ms

K = 1 on one H200 vs. K = 7 on eight H200s, per forward.

SCALING · WAN 720p / K = 216432.8 → 899.7 ms

One H200 → eight H200s, per forward. More GPUs do not always help.

Tripling Wan’s current-chunk latents (K = 1 → 3) raises single-GPU forward time by 2.05× at 480p and 2.74× at 720p. These profiles exclude text encoding, VAE decoding, and the sampler loop; they are not end-to-end response times.

Wan quality · all resolutions and step budgets

Separately trained K = 1/3/7/21 configurations; Toy6K training, 946 VBench prompts × five seeds, 81 frames at 16 fps. Total is the VBench aggregate; the remaining columns retain the reported component and auxiliary scores.

Wan1.3B quality measurements
ResolutionStepsKTotalQualitySemanticDynamicVisionRewardInstruction
480p1179.1181.8068.344.006.5090.00
480p1380.3883.4068.2958.004.8194.00
480p1781.3583.9770.8648.005.79100.00
480p12181.6884.8369.0885.009.9394.00
480p2180.0782.7669.3414.009.3198.00
480p2381.6284.6369.5757.009.82100.00
480p2780.8484.0667.9493.009.29100.00
480p22181.8484.6270.7471.0010.17100.00
480p4179.0081.3869.5013.007.6596.00
480p4381.6684.7069.5180.009.64100.00
480p4781.7084.7369.5678.0010.1998.00
480p42182.1385.0370.5485.009.6698.00
720p1178.9481.3769.194.005.8292.00
720p1379.9982.7069.1847.003.9398.00
720p1780.7883.0571.6741.005.2396.00
720p12180.4983.1070.0467.009.56100.00
720p2179.8682.4069.709.009.4398.00
720p2381.5284.1171.1848.0010.10100.00
720p2780.9783.7969.6784.009.2594.00
720p22180.2682.6170.8439.009.9198.00
720p4178.8481.1969.4422.007.1892.00
720p4381.5984.3670.4974.009.59100.00
720p4781.8684.6270.8373.009.7796.00
720p42181.6584.4370.5266.009.3696.00
Wan efficiency · H200 forward profiles

Batch one, BF16, a shared reference checkpoint, and synthetic inputs with a 21-latent prefix. Slowest-rank median over 30 synchronized forwards after 10 warm-ups, including communication and output gathering. USP1/2/4/8 use 1/2/4/8 GPUs; all values are milliseconds.

Wan H200 single-forward latency (ms)
ResolutionKUSP1USP2USP4USP8
480p177.563.765.972.1
480p3159.291.567.072.9
480p7358.3194.7105.278.6
480p211406.2750.0384.1211.0
720p1218.7143.495.872.7
720p3599.6318.4166.5119.4
720p71521.4809.2443.8276.4
720p216432.83335.51737.1899.7

LTX-2.3 shows why the choice is model-specific. Its six visual scores move differently with K, and small chunks do not benefit from more GPUs in these cached-forward profiles.

LTX-2.3 · visual quality and cached-forward profiles

Internal32: 32 matched cases/seeds, 361 frames at 832 × 448 and 24 fps; Internal8K training, one training run per K. Three DiT forwards with schedule [999, 757, 522, 0] (0 is the terminal boundary). Visual scores are ×100 and have no combined total. SC/BC: subject/background consistency; Aesth.: aesthetic quality; Imaging: imaging quality; Smooth.: motion smoothness.

LTX-2.3 visual scores (×100)
KSCBCAesth.ImagingSmooth.Dynamic
187.9794.6852.8273.0999.6231.25
388.8194.3152.2165.9899.5346.88
989.4293.2750.7268.9699.5634.38
1584.9692.2151.0165.9899.5078.12
4588.7293.6852.4068.7599.6440.62
LTX-2.3 H200 cached-forward latency (ms)
KUSP1USP2USP4USP8
1274.0440.5479.6526.1
3255.6423.5457.0511.2
9426.7417.8456.6502.0
15660.3420.4461.0509.4
452016.21158.4634.0507.9

Keep chunk size available as a quality–efficiency choice, while making commitment fine-grained.

Make joint generation faster.
Keep interaction fine-grained.

Decoupling generation and commitment opens another route to responsiveness: accelerate high-quality joint generation.

It also motivates exploring bidirectional teacher models for interactive streaming.