arXiv:2607.13031cs.LGcs.CV2026-07被引 1

视频扩散模型在长链条因果推理任务中表现下降,因缺乏可扩展的串行计算能力。

The Seriality Gap in Video Diffusion Models

论文配图:The Seriality Gap in Video Diffusion Models
图 1 · 摘自论文原文
  • 通过多球碰撞实验发现,序列事件越长,模型性能越差
  • 增加去噪步骤无法缓解性能下降,说明问题不在时间长度
  • 自回归生成和深层架构更有效,适合需要串行推理的任务

当一个球撞击另一个球并依次传递时,视频模型应预测每一步的后果。在多球刚性球动力学的控制实验中,标准双向视频扩散模型在因果链变长时性能下降,即使提供更多去噪步骤也无改善。在仅含单球、无球间相互作用的对照实验中,性能下降几乎消失,表明问题根源是事件间的依赖结构而非视频长度。干预研究显示,提升有效串行计算的方法(如自回归/分块生成、模型深度)能显著改善性能。我们提出‘串行性差距’:任务所需的串行计算量增长与扩散模型去噪环无法提供可扩展串行计算之间的不匹配。进一步证明,在确定性视频预测中,去噪步骤不增加超越主干网络的串行计算,揭示了视频扩散模型在串行推理与模拟任务中的结构性瓶颈。

原文摘要 · Abstract (English)

When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computation improve performance disproportionately, including autoregressive/blockwise generation and architectural depth. We identify this pattern as the seriality gap: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for deterministic video prediction, denoising steps do not add serial computation beyond the backbone, indicating a structural obstacle for video diffusion on serial reasoning and simulation tasks.

视频生成扩散模型串行推理因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。