用预训练视角模型生成360度沉浸视频,解决空间不连贯问题
ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
- 提出视点图表示法,兼顾全局连续性与细节
- 引入全景-视角注意力机制,利用预训练视角先验
- 实现动态一致的全景视频生成,性能领先
全景视频生成旨在合成360度沉浸式视频,在虚拟现实、世界模型和空间智能领域具有重要意义。现有方法因全景数据与主流训练数据(视角数据)之间存在固有模态差异,难以生成高质量全景视频。本文提出一种新框架,利用预训练视角视频模型生成全景视频。具体而言,设计了一种名为视点图(ViewPoint map)的新全景表示,同时具备全局空间连续性和精细视觉细节。通过提出的全景-视角注意力机制,模型能有效利用预训练视角先验,并捕捉视点图中的全景空间相关性。大量实验表明,该方法可生成高度动态且空间一致的全景视频,达到当前最佳性能,优于以往方法。
原文摘要 · Abstract (English)
Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspective data, which constitutes the majority of the training data for modern diffusion models. In this paper, we propose a novel framework utilizing pretrained perspective video models for generating panoramic videos. Specifically, we design a novel panorama representation named ViewPoint map, which possesses global spatial continuity and fine-grained visual details simultaneously. With our proposed Pano-Perspective attention mechanism, the model benefits from pretrained perspective priors and captures the panoramic spatial correlations of the ViewPoint map effectively. Extensive experiments demonstrate that our method can synthesize highly dynamic and spatially consistent panoramic videos, achieving state-of-the-art performance and surpassing previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。