用球面对极约束实现文本驱动全景图的高质量生成
DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion
- 基于球面对极约束设计多视角扩散模型,提升全景图一致性
- 在百万级全景视频-文本数据集上微调模型,支持多样生成
- 适合需要真实感全景内容的影视、游戏与虚拟现实开发者
基于扩散的方法在2D图像和3D物体生成中已取得显著进展,但3D场景乃至$360^{\ extdegree}$图像的生成仍受限于场景数据集规模有限、3D场景本身复杂以及多视角图像生成的一致性难题。为此,我们首先构建了一个包含数百万连续全景关键帧的大型全景视频-文本数据集,包含对应的全景深度、相机位姿和文本描述。随后提出一种新型文本驱动的全景生成框架DiffPano,实现可扩展、一致且多样的全景场景生成。具体而言,利用Stable Diffusion的强大生成能力,我们在构建的全景视频-文本数据集上通过LoRA微调单视图文本到全景扩散模型,并进一步设计球面对极感知的多视图扩散模型,以保证生成全景图像的多视角一致性。大量实验表明,DiffPano可在给定未见文本描述和相机位姿的情况下,生成可扩展、一致且多样的全景图像。
原文摘要 · Abstract (English)
Diffusion-based methods have achieved remarkable achievements in 2D image or 3D object generation, however, the generation of 3D scenes and even $360^{\circ}$ images remains constrained, due to the limited number of scene datasets, the complexity of 3D scenes themselves, and the difficulty of generating consistent multi-view images. To address these issues, we first establish a large-scale panoramic video-text dataset containing millions of consecutive panoramic keyframes with corresponding panoramic depths, camera poses, and text descriptions. Then, we propose a novel text-driven panoramic generation framework, termed DiffPano, to achieve scalable, consistent, and diverse panoramic scene generation. Specifically, benefiting from the powerful generative capabilities of stable diffusion, we fine-tune a single-view text-to-panorama diffusion model with LoRA on the established panoramic video-text dataset. We further design a spherical epipolar-aware multi-view diffusion model to ensure the multi-view consistency of the generated panoramic images. Extensive experiments demonstrate that DiffPano can generate scalable, consistent, and diverse panoramic images with given unseen text descriptions and camera poses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。