arXiv:2512.01481cs.CV2025-12被引 3

无需训练即可生成高保真多视角同步视频,解决4D世界建模难题

ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling

  • 构建时空约束的全局超空间表示世界状态
  • 通过超空间引导采样实现多视角扩散过程同步
  • 适用于需要实时多视角一致生成的虚拟场景应用

尽管当前基于相机控制的视频生成模型可产出电影级效果,但直接扩展至生成3D一致且高保真的时间同步多视角视频仍具挑战性,而这正是构建4D世界的关键能力。部分工作依赖数据增强或测试时优化,但受限于模型泛化能力弱与可扩展性差。为此,我们提出ChronosObserver,一种无需训练的方法,包含用于表征4D世界时空约束的「世界状态超空间」,以及利用超空间引导多视角扩散采样轨迹同步的「超空间引导采样」机制。实验表明,该方法在不微调扩散模型的前提下,实现了高保真、3D一致的时间同步多视角视频生成。

原文摘要 · Abstract (English)

Although prevailing camera-controlled video generation models can produce cinematic results, lifting them directly to the generation of 3D-consistent and high-fidelity time-synchronized multi-view videos remains challenging, which is a pivotal capability for taming 4D worlds. Some works resort to data augmentation or test-time optimization, but these strategies are constrained by limited model generalization and scalability issues. To this end, we propose ChronosObserver, a training-free method including World State Hyperspace to represent the spatiotemporal constraints of a 4D world scene, and Hyperspace Guided Sampling to synchronize the diffusion sampling trajectories of multiple views using the hyperspace. Experimental results demonstrate that our method achieves high-fidelity and 3D-consistent time-synchronized multi-view videos generation without training or fine-tuning for diffusion models.

视频生成扩散模型4D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。