arXiv:2509.11092cs.CVcs.AI2025-09被引 3

用低秩适配让普通视频模型生成高质量全景视频

PanoLora: Bridging Perspective and Panoramic Video Generation with LoRA Adaptation

  • 将全景生成视为视角转换问题,用LoRA高效微调预训练模型
  • 仅需约1000张视频即可实现高质全景生成,视觉效果更优
  • 适合想低成本接入全景视频生成的研究者和开发者

生成高质量360°全景视频仍面临巨大挑战,源于全景与传统透视视图在投影方式上的根本差异。透视视频基于单一视角、有限视野,而全景内容需渲染完整环境,导致标准视频生成模型难以适配。现有方法常引入复杂架构或大规模训练,效率低下且效果不佳。受低秩适配(LoRA)在风格迁移中成功启发,我们提出将全景视频生成视为从透视视图到全景视图的适配问题。理论分析表明,当LoRA秩超过任务自由度时,可有效建模两种投影间的变换。我们的方法仅用约1000个视频即可高效微调预训练视频扩散模型,实现高质量全景生成。实验表明,该方法保持了正确的投影几何结构,在视觉质量、左右一致性及运动多样性上均优于先前最先进方法。

原文摘要 · Abstract (English)

Generating high-quality 360° panoramic videos remains a significant challenge due to the fundamental differences between panoramic and traditional perspective-view projections. While perspective videos rely on a single viewpoint with a limited field of view, panoramic content requires rendering the full surrounding environment, making it difficult for standard video generation models to adapt. Existing solutions often introduce complex architectures or large-scale training, leading to inefficiency and suboptimal results. Motivated by the success of Low-Rank Adaptation (LoRA) in style transfer tasks, we propose treating panoramic video generation as an adaptation problem from perspective views. Through theoretical analysis, we demonstrate that LoRA can effectively model the transformation between these projections when its rank exceeds the degrees of freedom in the task. Our approach efficiently fine-tunes a pretrained video diffusion model using only approximately 1,000 videos while achieving high-quality panoramic generation. Experimental results demonstrate that our method maintains proper projection geometry and surpasses previous state-of-the-art approaches in visual quality, left-right consistency, and motion diversity.

全景视频LoRA扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。