用线性投影加速全景图生成,无需训练即可实现更自然的图像融合。
Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation

- 将潜在特征融合建模为正则化最小二乘问题,用迭代求解器高效计算。
- 仅需较少视角即可生成稳定全景图,相比基线提升15.36倍推理速度。
- 适合追求高效、高质量无训练全景生成的开发者与研究者。
我们提出LF-MultiDiffusion,一种无需训练的全景图生成方法,将MultiDiffusion扩展至支持目标与参考图像空间间的线性投影。核心思想是将潜在特征聚合重新表述为正则化最小二乘问题,并在去噪循环中使用基于克雷洛夫子空间的迭代求解器高效求解。该方法实现了比现有无训练方法更密集、更自然的映射,显著提升生成稳定性,同时大幅减少所需视角数量。实验表明,LF-MultiDiffusion在视觉质量、文本对齐和全景一致性上优于最强的无训练基线,且推理效率提升15.36倍。
原文摘要 · Abstract (English)
We propose LF-MultiDiffusion, a training-free panorama generation method that extends MultiDiffusion to support linear projections between target and reference image spaces. Our key idea is to reformulate latent aggregation as a regularized least-squares problem and solve it efficiently with a Krylov-based iterative solver inside the denoising loop. This formulation enables denser and more natural mappings than prior training-free methods, yielding more stable generation with far fewer perspective views. As a result, LF-MultiDiffusion reduces the number of image generator evaluations during denoising and significantly improves inference efficiency. Experiments show that LF-MultiDiffusion achieves better visual quality, text alignment, and panoramic consistency than the strongest training-free baseline, while providing a 15.36$\times$ speedup. Our project page is available at: https://ahykw.github.io/lfmd.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。