生成可自由探索的全景视频世界,解决视角受限与相机控制差问题
PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
- 基于虚拟环境模拟全景路径,构建大规模视频探索数据集
- 提出球面感知扩散架构,提升全景视频的视觉质量和时序连续性
- 支持多样化相机轨迹,适合沉浸式应用与智能体自由探索
生成完整且可自由探索的360度视觉世界可支持多种下游应用。尽管已有研究取得进展,但仍受限于视场狭窄导致场景不连贯,或相机控制不足限制用户或自主代理的自由探索。为此,我们提出PanoWorld-X,一种高保真、可控的全景视频生成框架,支持多样相机轨迹。首先,通过Unreal Engine在虚拟3D环境中模拟相机路径,构建大规模全景视频-探索路径配对数据集。由于全景数据的球面几何与传统视频扩散模型的归纳偏置不匹配,我们引入球面感知扩散变换器(Sphere-Aware Diffusion Transformer),将等距柱状投影特征重新映射到球面,以在潜在空间中建模几何邻接关系,显著提升视觉保真度和时空连续性。大量实验表明,PanoWorld-X在运动范围、控制精度和视觉质量等方面均表现优异,展现出在真实场景中的应用潜力。
原文摘要 · Abstract (English)
Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient camera controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. Specifically, we first construct a large-scale dataset of panoramic video-exploration route pairs by simulating camera trajectories in virtual 3D environments via Unreal Engine. As the spherical geometry of panoramic data misaligns with the inductive priors from conventional video diffusion, we then introduce a Sphere-Aware Diffusion Transformer architecture that reprojects equirectangular features onto the spherical surface to model geometric adjacency in latent space, significantly enhancing visual fidelity and spatiotemporal continuity. Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。