๐ Highlights
DeCoPrune frames historical KV-cache compression for autoregressive video diffusion as a denoising-consistency problem: tokens whose intermediate clean prediction remains far from the final denoised value are considered more informative and are retained. This is a useful builder-facing idea because it converts a difficult long-range memory question into a model-intrinsic score available during normal generation.
The strongest reported result is on LingBot World v2 with CMBench: DeCoPrune reaches 0.6701 DINO at 85.43% PR and 4.14x continuation speedup, while DeCoPrune-HS reaches 0.6783 DINO at 86.19% PR, close to FullKV's 0.6803 DINO. The central caveat is that the paper's strongest minute-scale evidence is concentrated on one primary backbone and on a benchmark centered on localized visual recall.
- Training-free KV-cache pruning based on step-to-final denoising discrepancy.
- CMBench evaluates 116 Reappear/Revisit tasks across 58 approximately one-minute contexts.
- DeCoPrune keeps high-discrepancy tokens and prunes low-discrepancy tokens after a recent-window delay.
- Main comparison reports near-FullKV DINO with more than 85% cumulative historical-token pruning.
- Ablations show that reverse low-discrepancy retention performs substantially worse than DeCoPrune at matched PR.
๐ฏ Introduction
The paper targets autoregressive video diffusion, where generation proceeds chunk by chunk and each new chunk attends to historical KVs from the preceding video. This design naturally supports streaming generation and interactive control, but the cost of memory and attention grows with rollout length. Full history retention preserves evidence for long-range consistency, yet becomes increasingly impractical; aggressive truncation improves efficiency but risks removing the visual evidence needed to reproduce a previously seen object, person, or scene.
The authors argue that existing compression strategies are misaligned with the core question of contextual redundancy. Fixed-window methods can discard useful old events, and local temporal-difference strategies compare neighboring frames rather than deciding whether a chunk contributes information beyond the retained context. DeCoPrune's objective is therefore to score each current chunk against the retained cache itself, so repeated or context-predictable content can be pruned while unpredictable visual evidence remains available to future continuations.
The paper also introduces CMBench because general video quality metrics may not expose failures of historical recall. CMBench asks the model to continue from an approximately one-minute context and reproduce or revisit specific earlier targets. That makes the benchmark closer to the operational failure mode builders care about in long-context generation: the output can look plausible while still forgetting the exact item or view that the user expected the system to preserve.
๐ฌ Methodology
The core hypothesis is that tokens already predictable from the retained cache tend to reach stable clean predictions earlier in the denoising trajectory, whereas tokens containing information not explained by the cache require greater refinement. During normal generation of chunk i, DeCoPrune records the clean prediction at probe timestep tau* and later compares it with the final denoised chunk from the same trajectory. The resulting token-wise discrepancy becomes a retention criterion: high-discrepancy tokens enter the long-term cache, and low-discrepancy tokens are treated as redundant.
The scoring rule is explicit: each token p receives a mean squared discrepancy between the intermediate clean prediction and final denoised value, then a threshold gamma converts it into a binary mask. The paper uses this as a shared mask across layers, so the retained historical KVs are gathered physically rather than merely ignored by an attention mask. The system protects the initial sink chunk and keeps the most recent w chunks dense, storing masks until older chunks leave the recent window. This delay preserves local evidence for immediate successors while still reducing long-term history.
For observed video contexts, the original denoising trajectory is unavailable, so the method re-noises each finalized context chunk to tau*, evaluates the context-conditioned clean prediction at the original absolute temporal position, and compares it with the observed clean chunk. The implementation further adds RoPE re-indexing, which compresses temporal coordinates of older keys into a fixed virtual span while preserving recent positions. The head-specialized variant DeCoPrune-HS uses the attention-based partition from ForcingKV: dynamic heads apply DeCoPrune, while static heads use Streaming.
Figure fig:overview captures the builder-critical pipeline: probe prediction and final prediction produce a discrepancy map, the threshold produces a retention mask, and the historical KV cache is compacted for future chunks. The method remains training-free and keeps the generator frozen, which lowers adoption cost if the target video diffusion system exposes intermediate clean predictions and supports physical KV gathering.
equation 1
Meaning: Controls token scoring and retention; gamma sets how aggressively low-discrepancy tokens are pruned.
equation 2
Meaning: Controls the physical cache update by gathering only retained K and V entries for each layer.
๐ Experiments
The main empirical setup uses LingBot World v2 as the primary backbone and runs experiments on four NVIDIA H200 GPUs. CMBench is the primary benchmark: it contains 58 approximately one-minute contexts and 116 continuation tasks, with 50 H3-generated episodes providing 102 tasks and eight real-world episodes providing 14 tasks. The two task types are Reappear, where an observed person or object must appear again, and Revisit, where the continuation must return to a previously observed scene with a salient target object.
The evaluation localizes the target in generated frames with OWL-ViT, segments reference and generated targets with SAM 2, and computes DINOv2 embedding similarity over crops. The task score is the maximum cosine similarity across continuation frames, allowing the requested event to occur at any point; the score is zero if the target is never detected. PR measures cumulative historical KV-token reduction across generated chunks, layers, and heads, excluding the current noisy chunk. FPS and speedup measure continuation generation only and exclude prefix processing.
The main comparison includes FullKV, Streaming, DummyForcing, ForcingKV, TempDiff, Random in the selection ablation, DeCoPrune, and DeCoPrune-HS. Table tab:cmbench_results reports that FullKV is the uncompressed reference at 0.6803 DINO and 1.568 FPS. DeCoPrune reaches 0.6701 DINO with 85.43% PR, 6.489 FPS, and 4.14x speedup. DeCoPrune-HS reaches 0.6783 DINO at 86.19% PR and 4.09x speedup, leaving only a 0.0020 DINO gap to FullKV in the main comparison. TempDiff reaches 0.6229 DINO at 86.36% PR, so DeCoPrune improves DINO by 0.0472 at a similar compression level.
The ablations test whether the result comes from the consistency signal rather than only cache budget or positional correction. RoPE re-indexing improves DINO for methods retaining distant content, including FullKV, DeCoPrune-HS, TempDiff, and Random, but has little or negative effect on Streaming, DummyForcing, and ForcingKV. Token-selection direction is more decisive: at nearly matched PR, DeCoPrune reaches 0.6701 DINO, Random reaches 0.6091, and the matched-PR reverse criterion reaches 0.4512. Even the reverse criterion with PR 18.79% reaches only 0.5621, supporting the authors' claim that high-discrepancy retention is the useful direction.
tab:cmbench_results
Caption: Main comparison on CMBench and VBench. DINO is reported on a 0--1 scale; results are averaged over three random seeds. FPS and speedup over FullKV measure continuation generation only, excluding prefix processing.
| Method | DINO โ | PR โ | FPS โ | Speedup โ | Temporal Flickering โ | Motion Smoothness โ | Aesthetic Quality โ | Image Quality โ |
|---|---|---|---|---|---|---|---|---|
| FullKV | 0.6803 | 0.00% | 1.568 | 1.00ร | 0.9499 | 0.9705 | 0.4633 | 0.7125 |
| Streaming (Xu et al., 2026c; Yang et al., 2025) | 0.4592 | 94.37% | 9.591 | 6.12ร | 0.9474 | 0.9698 | 0.4769 | 0.7103 |
| DummyForcing (Guo et al., 2026) | 0.4461 | 98.03% | 7.393 | 4.72ร | 0.9425 | 0.9687 | 0.4677 | 0.6933 |
| ForcingKV (Ji et al., 2026) | 0.5313 | 80.62% | 5.288 | 3.37ร | 0.9437 | 0.9661 | 0.4610 | 0.7018 |
| TempDiff (Hwang et al., 2024; Fu et al., 2025) | 0.6229 | 86.36% | 6.443 | 4.11ร | 0.9425 | 0.9634 | 0.4677 | 0.7161 |
| DeCoPrune (ours) | 0.6701 | 85.43% | 6.489 | 4.14ร | 0.9424 | 0.9650 | 0.4736 | 0.7067 |
| DeCoPrune-HS (ours) | 0.6783 | 86.19% | 6.414 | 4.09ร | 0.9428 | 0.9652 | 0.4634 | 0.7019 |
Why It Matters: It reports the paper's main speed, pruning, recall, and VBench comparison.
tab:reindex_ablation
Caption: RoPE re-indexing ablation. โDINO is the change from re-indexing.
| Method | w/ re-index DINO โ | w/o re-index DINO โ | โDINO |
|---|---|---|---|
| FullKV | 0.6804 | 0.4385 | +0.2418 |
| DeCoPrune-HS (ours) | 0.7156 | 0.6686 | +0.0470 |
| TempDiff | 0.6365 | 0.6019 | +0.0346 |
| Random | 0.6723 | 0.5515 | +0.1208 |
| DummyForcing | 0.4005 | 0.4012 | -0.0007 |
| ForcingKV | 0.4825 | 0.4860 | -0.0035 |
| Streaming | 0.3994 | 0.4040 | -0.0046 |
Why It Matters: It isolates the role of positional correction in retrieving retained long-range content.
tab:selection_direction_ablation
Caption: Token-selection ablation. Reverse retains low-discrepancy tokens.
| Policy | DINO โ | PR โ |
|---|---|---|
| DeCoPrune | 0.6701 | 85.43% |
| Random | 0.6091 | 85.31% |
| Reverse criterion | 0.5621 | 18.79% |
| Reverse criterion (matched PR) | 0.4512 | 85.39% |
Why It Matters: It validates the direction of the denoising-consistency criterion.
equation 3
Meaning: Defines how CMBench scores visual recall over continuation frames.
equation 4
Meaning: Defines cumulative historical KV-token reduction relative to FullKV.
fig:overview

Caption: DeCoPrune overview. A probe and final prediction from the same denoising trajectory yield a token-retention mask. The mask physically compacts historical KVs after the recent-window delay.
Why It Matters: This is the clearest architecture figure for builders: it shows how the probe prediction, final prediction, discrepancy score, retention mask, and KV compaction connect in the online cache update.
fig:cmbench_overview

Caption: Overview of CMBench. Top: a one-minute generated context assembled from six prompted clips; three target events yield three continuation tasks. Bottom: the corresponding continuation prompts and evaluation pipeline. The reference target and its generated counterpart are localized and segmented, then compared using DINO similarity.
Why It Matters: This grounds the benchmark design, showing that the evaluation is about recalling specific prior visual evidence rather than producing a plausible generic continuation.
fig:qualitative_comparison

Caption: Qualitative comparison under the matched settings used in Tables~tab:cmbench_results and~tab:selection_direction_ablation. We highly recommend viewing the video comparisons on the supplementary webpage to better appreciate temporal consistency and visual detail.
Why It Matters: The qualitative comparisons complement the aggregate DINO results by showing case-level behavior under the same settings as the main comparison and selection ablation.
fig:probe_threshold_selection

Caption: Ablation and metric analysis on 13 independent cases. (a) Threshold $ $ is swept from light to dark at each probe-step index; the marked operating point is $s^*=2$ ($ ^*=899$), $ =0.10$. (b) VBench subject/background consistency and CMBench DINO versus PR.
Why It Matters: This supports the chosen probe setting and shows why CMBench is more sensitive than saturated subject/background consistency metrics for recall-preserving compression.
๐ฎ Conclusion
The paper's main conclusion is that denoising consistency is a practical proxy for token importance in long-context autoregressive video diffusion. DeCoPrune uses that proxy to reduce the historical KV cache while preserving much of the context recall measured by CMBench. The result is strongest as a training-free systems method: it keeps the generator frozen, reuses normal denoising predictions, and reports a large reduction in cumulative historical KV tokens with near-FullKV DINO on the primary setup.
For builders, the takeaway is not just that pruning can be aggressive, but that recall-aware pruning needs a metric and selection signal aligned with memory. The paper shows that generic visual quality metrics may stay stable even when contextual recall deteriorates, so deployment evaluation should include continuation tasks that verify exact prior evidence.
๐ ๏ธ Future Research Improvements
The main-body evidence suggests several next steps, while leaving some questions open. First, the strongest minute-scale study should be broadened across additional autoregressive video diffusion backbones rather than relying primarily on LingBot World v2. The paper mentions supplementary comparisons for other backbones, but the main body's central numeric story is tied to LingBot World v2 and CMBench.
Second, CMBench could be extended beyond localized object/person/view recall. The current protocol is valuable because it is reference-grounded and measurable, but long video memory also includes action continuity, causal state, spatial layout, and multi-event relationships. Third, system profiling should distinguish cumulative historical token reduction, peak memory, prefix-processing cost, and end-to-end latency, because the reported PR and continuation FPS do not fully characterize deployment resource usage.
- Broaden main-body backbone coverage for minute-scale contexts.
- Add recall tasks that go beyond localized visual target matching.
- Report peak memory and end-to-end latency alongside cumulative PR and continuation FPS.
- Study automatic threshold selection beyond the reported gamma=0.10 operating point.
๐ญ Potential Industry Use Scenarios
DeCoPrune is most relevant to products that need long-running or interactive video generation without retaining an ever-growing full KV history. Examples include first-person world simulators, interactive creative tools, virtual production previews, game-like scene exploration, and agentic video systems where a model must return to earlier objects or views after intervening motion. In those settings, users often notice failures as identity drift or missing previously established props, not merely as lower frame quality.
The method is also plausible as an infrastructure optimization for hosted video generation because it is training-free and uses intrinsic signals from the generation trajectory. A service could potentially offer longer continuations or lower continuation latency if the underlying model exposes suitable intermediate predictions and cache manipulation hooks. The evidence does not prove readiness across all production systems, however: the main-body metrics report continuation speed only, exclude prefix processing, and evaluate recall through CMBench's DINO-based target protocol rather than product-specific user studies.
- Interactive video generation where earlier objects must reappear consistently.
- Long-rollout world models that need bounded KV-cache growth.
- Creative editing workflows that revisit prior scenes or props.
- Hosted inference stacks seeking continuation speedups without retraining.
๐ฌ Critical Analysis
The paper's strongest technical move is aligning the pruning signal with the actual conditioning context. Instead of assuming that old tokens are useless or that local frame differences capture importance, DeCoPrune asks whether a token's final denoised value was already predictable from the retained context at an intermediate step. The Table 3 ablation is especially persuasive because the reverse criterion performs poorly: at nearly matched PR, DeCoPrune reaches 0.6701 DINO while reverse selection reaches 0.4512, and even the low-compression reverse setting reaches only 0.5621.
The benchmark contribution is also useful, because CMBench directly probes memory through Reappear and Revisit tasks. That said, the evaluation is only as broad as the target localization and DINO similarity protocol. DINO crop similarity is a practical signal for visual identity recall, but it may miss temporal ordering, physical plausibility, or semantic consistency that does not reduce to a localized crop. The main body responsibly reports VBench dimensions, yet those metrics remain broadly comparable across policies and therefore do not substitute for recall-specific evaluation.
Reproducibility is reasonably supported in the main body: the paper states the backbone, H200 hardware, noise schedule, probe timestep, threshold, default video configuration, baselines, and cache-budget details for the main comparison. The remaining builder risks are integration-specific. A model must expose intermediate clean predictions, allow per-token KV gathering across layers and heads, and tolerate RoPE re-indexing. The reported PR is cumulative rather than peak memory, and the speedup excludes prefix processing, so production gains should be remeasured under full serving conditions.