无需相机参数,直接将普通图像视频转为360全景图
360Anything: Geometry-Free Lifting of Images and Videos to 360°
- 用扩散模型把透视图当序列处理,不依赖几何对齐
- 在图像和视频转换上超越已有方法,连用真实相机数据的模型也比不过
- 解决全景拼接缝问题,还能零样本估计视角和方向
将透视图像和视频升维至360°全景图,可实现沉浸式三维世界生成。现有方法通常依赖透视图与等距矩形投影(ERP)空间间的显式几何对齐,但需已知相机参数,在真实场景数据中常因参数缺失或噪声而受限。本文提出360Anything,一种基于预训练扩散变压器的无几何框架。通过将透视输入与全景目标视为纯令牌序列,360Anything以数据驱动方式学习透视到ERP的映射,无需相机信息。该方法在图像与视频的透视到360°生成任务上均达当前最优性能,超越依赖真实相机信息的先前工作。我们还揭示了ERP边界伪影的根本原因在于VAE编码器中的零填充,并提出环形潜在编码以实现无缝生成。最后,我们在零样本相机视场角与姿态估计基准上取得竞争力结果,表明360Anything具备深层几何理解能力,具更广泛计算机视觉应用潜力。更多结果见https://360anything.github.io/。
原文摘要 · Abstract (English)
Lifting perspective images and videos to 360° panoramas enables immersive 3D world generation. Existing approaches often rely on explicit geometric alignment between the perspective and the equirectangular projection (ERP) space. Yet, this requires known camera metadata, obscuring the application to in-the-wild data where such calibration is typically absent or noisy. We propose 360Anything, a geometry-free framework built upon pre-trained diffusion transformers. By treating the perspective input and the panorama target simply as token sequences, 360Anything learns the perspective-to-equirectangular mapping in a purely data-driven way, eliminating the need for camera information. Our approach achieves state-of-the-art performance on both image and video perspective-to-360° generation, outperforming prior works that use ground-truth camera information. We also trace the root cause of the seam artifacts at ERP boundaries to zero-padding in the VAE encoder, and introduce Circular Latent Encoding to facilitate seamless generation. Finally, we show competitive results in zero-shot camera FoV and orientation estimation benchmarks, demonstrating 360Anything's deep geometric understanding and broader utility in computer vision tasks. Additional results are available at https://360anything.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。