用文本或单视角视频生成360度沉浸式视频
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
- 通过多视角注意力增强扩散模型,实现多视角视频一致性生成
- 支持长视频自回归生成,训练时随机采样帧与视角提升效率
- 在真实与合成数据集上均优于现有方法,生成更逼真全景视频
高分辨率全景视频对虚拟现实中的沉浸体验至关重要,但其采集需特殊设备和复杂相机布置,极为困难。本文提出VideoPanda,一种基于文本或单视角视频条件生成360°视频的新方法。VideoPanda通过引入多视角注意力层增强视频扩散模型,使其能够生成视角一致的多视图视频,并拼接为沉浸式全景内容。该模型联合训练于文本仅输入和单视角视频两种条件,支持长视频的自回归生成。为缓解多视角视频生成的计算负担,训练时随机子采样时间长度与相机视角,结果表明模型可在推理阶段优雅泛化至生成更多帧。在真实世界与合成视频数据集上的大量评估显示,相比现有方法,VideoPanda在所有输入条件下均生成了更真实、连贯的360°全景图像。项目网站:https://research.nvidia.com/labs/toronto-ai/VideoPanda/
原文摘要 · Abstract (English)
High resolution panoramic video content is paramount for immersive experiences in Virtual Reality, but is non-trivial to collect as it requires specialized equipment and intricate camera setups. In this work, we introduce VideoPanda, a novel approach for synthesizing 360$^\circ$ videos conditioned on text or single-view video data. VideoPanda leverages multi-view attention layers to augment a video diffusion model, enabling it to generate consistent multi-view videos that can be combined into immersive panoramic content. VideoPanda is trained jointly using two conditions: text-only and single-view video, and supports autoregressive generation of long-videos. To overcome the computational burden of multi-view video generation, we randomly subsample the duration and camera views used during training and show that the model is able to gracefully generalize to generating more frames during inference. Extensive evaluations on both real-world and synthetic video datasets demonstrate that VideoPanda generates more realistic and coherent 360$^\circ$ panoramas across all input conditions compared to existing methods. Visit the project website at https://research.nvidia.com/labs/toronto-ai/VideoPanda/ for results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。