arXiv:2605.25449cs.CV2026-05中稿 · CVPR被引 1

用360°视频扩散模型生成数字孪生,解决视角不一致问题

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion

论文配图:Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
图 1 · 摘自论文原文
  • 通过3D缓存构建几何骨架,支撑任意相机路径
  • 生成视频视觉质量高,几何一致性显著优于现有方法
  • 适合需要真实场景模拟的数字孪生应用

从视频生成完整数字孪生需精确相机控制、全局场景覆盖及严格的时空一致性,但透视视频生成因视场有限,难以满足要求。窄视场导致长或多视角轨迹,加剧跨视角不一致与时间漂移。我们提出360°视频生成是天然解决方案:全景覆盖简化轨迹设计,提供强全局上下文以维持连贯性。引入Pantheon360:一种可控的360°视频生成框架,从稀疏360°输入合成高保真视频。核心思想是显式构建3D缓存,由输入重建,作为任意用户定义相机路径的几何骨架。使扩散模型专注于纹理细节优化,而3D缓存保障全局几何一致性。实验表明,Pantheon360在视觉质量和几何一致性上均表现优异,支持可靠灵活的360°场景生成,适用于下游仿真与数字孪生应用。

原文摘要 · Abstract (English)

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view trajectories, amplifying cross-view inconsistency and temporal drift. We argue that 360° video generation offers a natural solution: panoramic coverage simplifies trajectory design and provides a strong global context for maintaining coherence. We introduce Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion, a controllable 360° video generation framework that synthesizes high-fidelity videos from sparse 360° inputs. The key idea is an explicit 3D Cache, reconstructed from the input, which serves as a geometric scaffold for any user-defined camera path. This allows the diffusion model to focus on photorealistic texture refinement while the 3D Cache enforces global geometric consistency. Experiments show that Pantheon360 achieves superior visual quality and unmatched geometric coherence, enabling reliable and flexible 360° scene generation for downstream simulation and digital-twin applications.

数字孪生360°视频扩散模型3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。