arXiv:2607.05376cs.CVcs.GR2026-07中稿 · ECCV

实现任意长度与视角的动态场景多视角视频生成。

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

论文配图:MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
图 1 · 摘自论文原文
  • 用4D几何桥连接逐视图生成,统一时序与视角自回归。
  • 支持无固定时长限制的生成,单模型可产任意长度视频。
  • 适合需要高几何一致性的3D视频生成研究者。

近期视频扩散模型可实现长时单视角生成或短时多视角合成,但动态场景的长时多视角一致性生成仍无解。本文提出MV-Forcing框架,通过引入4D几何桥,在单一扩散模型中融合时序与视角自回归。核心思路是:基于已完成的源视角,重建其3D结构并渲染下一目标视角的几何先验,再由扩散模型优化为高质量视频。为突破教师模型固定时长限制,引入联合去噪机制,训练时双视角均从噪声初始化,实现无界时长生成。通过时空自强制分布匹配蒸馏,消除训练-推理暴露偏差。在合成与真实数据上的大量实验表明,该方法仅用一个少步学生模型即可生成任意长度与视角数量的几何一致多视角动态视频。

原文摘要 · Abstract (English)

Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.

视频生成多视角扩散模型4D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。