arXiv:2608.19556cs.CVcs.AI2026-08

Stream4D让流式视频生成保持动态一致性,避免画面僵化。

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

论文配图:Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
图 1 · 摘自论文原文
  • 用4D动态重建替代静态判别器,捕捉场景运动变化。
  • 在多个模型和生成时长下,显著提升运动连贯性和视觉质量。
  • 适合追求真实动态视频生成的研究者与开发者。

流式自回归扩散视频模型可实现实时、长时程视频生成,但其训练目标侧重局部帧预测,而非保持世界几何与动态的一致性:长期生成会积累几何漂移,导致画面变静或运动不自然。现有双向方法通过3D高斯溅射重建构建奖励信号来缓解此问题,但单一刚性3D重建无法建模动态场景,将真实物体运动误判为重建误差,反而鼓励视频冻结。该捷径在自回归设定中尤为致命,因每一段生成都会传播已僵化的状态。本文提出Stream4D,以前馈式4D重建奖励取代静态判别器,显式建模场景动态,使连贯运动获得高一致性奖励;同时引入运动先验,奖励合理运动幅值,惩罚抖动与非刚性伪影。最终方案结合两项奖励与轻量级感知锚点,在多种自回归视频骨干网络与生成时长下,显著提升4D重建质量,更有效保留运动特征,并获得更高人类偏好评分。

原文摘要 · Abstract (English)

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/

视频生成扩散模型动态建模自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。