arXiv:2510.10670cs.CV2025-10被引 3

用视频生成模型自动规划4D场景的拍摄视角,效果优于现有方法。

AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes

  • 分两阶段适配预训练视频模型,注入4D场景信息并生成带视角的视频。
  • 通过扩散模型从生成视频中反推相机外参,实现视角精准提取。
  • 适用于需要自然视角生成的4D场景交互任务,如虚拟拍摄与重建。

近期文本到视频(T2V)模型展现出强大的真实世界几何与物理规律视觉模拟能力,表明其具备作为隐式世界模型的潜力。受此启发,我们探索利用视频生成先验来实现从给定4D场景中进行视角规划的可行性,因为视频天然伴随动态场景中的自然视角。为此,我们提出一种兼容性两阶段范式,将预训练的T2V模型适配用于视角预测。首先,通过自适应学习分支将视角无关的4D场景表示注入预训练的T2V模型,生成的条件视频在视觉上嵌入了对应视角。随后,将视角提取建模为混合条件引导的相机外参去噪过程,具体是在预训练的T2V模型上引入相机外参扩散分支,以生成视频和4D场景为输入。实验结果表明,所提方法在多个指标上优于现有竞争者,消融研究验证了关键设计的有效性。在一定程度上,本工作证明了视频生成模型在真实世界4D交互中的潜力。

原文摘要 · Abstract (English)

Recent Text-to-Video (T2V) models have demonstrated powerful capability in visual simulation of real-world geometry and physical laws, indicating its potential as implicit world models. Inspired by this, we explore the feasibility of leveraging the video generation prior for viewpoint planning from given 4D scenes, since videos internally accompany dynamic scenes with natural viewpoints. To this end, we propose a two-stage paradigm to adapt pre-trained T2V models for viewpoint prediction, in a compatible manner. First, we inject the 4D scene representation into the pre-trained T2V model via an adaptive learning branch, where the 4D scene is viewpoint-agnostic and the conditional generated video embeds the viewpoints visually. Then, we formulate viewpoint extraction as a hybrid-condition guided camera extrinsic denoising process. Specifically, a camera extrinsic diffusion branch is further introduced onto the pre-trained T2V model, by taking the generated video and 4D scene as input. Experimental results show the superiority of our proposed method over existing competitors, and ablation studies validate the effectiveness of our key technical designs. To some extent, this work proves the potential of video generation models toward 4D interaction in real world.

视频生成4D场景视角规划扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。