arXiv:2505.22016cs.CV2025-05NeurIPS被引 19

将传统视频生成模型迁移到全景视频,解决畸变与边界衔接问题。

PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms

  • 引入纬度感知采样,避免全景图纬向畸变。
  • 在公开数据集上实现顶尖生成质量,支持零样本下游任务。
  • 构建高质量全景视频数据集PanoVid,覆盖多样场景。

全景视频生成可实现沉浸式360°内容创作,在需场景一致世界探索的应用中具有价值。然而,现有全景视频生成模型难以利用传统文本到视频模型的预训练生成先验,受限于数据规模不足及空间特征表示差异。本文提出PanoWan,通过极小模块有效将预训练文本到视频模型迁移至全景领域。PanoWan采用纬度感知采样以避免纬向畸变,结合旋转语义去噪与填充像素级解码,确保经度边界处无缝过渡。为支撑这些迁移表征的学习,我们构建了高质全景视频数据集PanoVid,包含带注释的多样化场景。实验表明,PanoWan在全景视频生成上达到当前最优性能,并展现出对零样本下游任务的强大鲁棒性。

原文摘要 · Abstract (English)

Panoramic video generation enables immersive 360° content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks. Our project page is available at https://panowan.variantconst.com.

全景视频扩散模型生成迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。