arXiv:2605.15391cs.CVcs.AI2026-05

让全景视频生成更符合三维空间逻辑,解决画面扭曲与运动不连贯问题。

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

论文配图:PanoWorld: Geometry-Consistent Panoramic Video World Modeling
图 1 · 摘自论文原文
  • 将全景视频生成视为几何与动态一致的潜在状态建模问题
  • 在多个数据集上显著提升深度一致性与轨迹连续性,视觉质量仍具竞争力
  • 适合需要真实空间理解的智能体应用,如机器人导航与虚拟现实

我们提出PanoWorld,一种从单张图像和文本描述生成几何一致的360°全景视频的世界模型。现有方法主要优化视觉真实感,未显式约束底层三维场景状态,导致输出虽看似合理却存在深度不一致、对应关系断裂及球面运动不合理等问题。为此,我们将全景视频生成重新定义为几何与动态一致的潜在状态建模任务。基于预训练的视角视频世界模型,引入两个轻量级正则化项:针对伪真值全景深度的深度一致性损失,以及监督追踪点在时间维度上三维世界帧位置的轨迹一致性损失。同时对条件输入和位置编码进行球面几何感知适配。我们还构建了PanoGeo,一个统一的几何感知全景视频数据集,涵盖真实与合成来源,具备一致的深度、轨迹与提示标注,用于训练和分层评估。实验表明,PanoWorld在几何一致性上优于已有全景生成方法,同时保持良好视觉真实感,证明全景视频生成需作为几何建模问题处理,以满足具身智能应用的空间理解需求。代码已开源。

原文摘要 · Abstract (English)

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry- and dynamics-consistent latent state modeling problem rather than pure visual synthesis. Building on a pre-trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground-truth panoramic depth, and a trajectory consistency loss that supervises the 3D world-frame positions of tracked points across time. We further apply spherical-geometry-aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry-aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications. Code is available at https://github.com/ostadabbas/PanoWorld.

全景视频三维建模世界模型几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。