arXiv:2507.07133cs.GRcs.AI2025-07被引 3

用扩散模型生成无缝全景图,解决多视角图像拼接难题

Generative Panoramic Image Stitching

  • 基于多参考图微调扩散模型,保持场景内容与布局一致
  • 在真实数据集上显著优于基线方法,无鬼影等伪影
  • 适合需要高质量全景生成的视觉创作与虚拟现实应用

我们提出生成式全景图像拼接任务,旨在合成与包含视差、光照差异、相机设置或风格变化的多张参考图像内容一致的无缝全景图。传统拼接方法在此复杂场景下会产生鬼影等伪影。尽管近期生成模型可实现多参考图的一致外推,但在合成大范围连贯区域时表现不佳。为此,我们提出一种方法:微调基于扩散的修复模型,使其依据多参考图像保留场景内容与布局。模型训练完成后,仅需单张参考图即可外推生成完整全景图,输出无缝且视觉一致的结果,忠实融合所有参考图像内容。在真实采集数据集上的评估表明,该方法在图像质量、结构一致性与场景布局保持方面显著优于基线。

原文摘要 · Abstract (English)

We introduce the task of generative panoramic image stitching, which aims to synthesize seamless panoramas that are faithful to the content of multiple reference images containing parallax effects and strong variations in lighting, camera capture settings, or style. In this challenging setting, traditional image stitching pipelines fail, producing outputs with ghosting and other artifacts. While recent generative models are capable of outpainting content consistent with multiple reference images, they fail when tasked with synthesizing large, coherent regions of a panorama. To address these limitations, we propose a method that fine-tunes a diffusion-based inpainting model to preserve a scene's content and layout based on multiple reference images. Once fine-tuned, the model outpaints a full panorama from a single reference image, producing a seamless and visually coherent result that faithfully integrates content from all reference images. Our approach significantly outperforms baselines for this task in terms of image quality and the consistency of image structure and scene layout when evaluated on captured datasets.

全景生成扩散模型图像拼接

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。