首个基于扩散模型的全景视频生成框架,实现精准相机轨迹控制
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
- 提出全景普吕克嵌入,通过球坐标转换编码相机位姿
- 设计球面基线模块,沿基线自适应掩码注意力提升跨视角一致性
- 在真实相机轨迹下生成高质量全景视频,显著优于现有方法
近期,相机可控视频生成发展迅速,实现了对视频生成更精确的控制。然而,现有方法主要聚焦于透视投影下的相机控制,而几何一致的全景视频生成仍具挑战性,主要源于全景姿态表示和球面投影的固有复杂性。为此,我们提出CamPVG,首个基于扩散模型、由精确相机姿态引导的全景视频生成框架。通过球面投影实现全景图像的相机位置编码,并引入跨视角特征聚合机制。具体地,提出全景普吕克嵌入,利用球坐标变换编码相机外参,有效捕捉全景几何结构,克服了传统方法在等距柱状投影下的局限性。同时,设计球面基线模块,通过沿基线的自适应注意力掩码施加几何约束,实现细粒度跨视角特征聚合,显著提升生成全景视频的质量与一致性。大量实验表明,该方法生成的全景视频与相机轨迹高度一致,远超现有方法在全景视频生成上的表现。
原文摘要 · Abstract (English)
Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camera control in perspective projection video generation, while geometrically consistent panoramic video generation remains challenging. This limitation is primarily due to the inherent complexities in panoramic pose representation and spherical projection. To address this issue, we propose CamPVG, the first diffusion-based framework for panoramic video generation guided by precise camera poses. We achieve camera position encoding for panoramic images and cross-view feature aggregation based on spherical projection. Specifically, we propose a panoramic Plücker embedding that encodes camera extrinsic parameters through spherical coordinate transformation. This pose encoder effectively captures panoramic geometry, overcoming the limitations of traditional methods when applied to equirectangular projections. Additionally, we introduce a spherical epipolar module that enforces geometric constraints through adaptive attention masking along epipolar lines. This module enables fine-grained cross-view feature aggregation, substantially enhancing the quality and consistency of generated panoramic videos. Extensive experiments demonstrate that our method generates high-quality panoramic videos consistent with camera trajectories, far surpassing existing methods in panoramic video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。