用检索技术提升跨视角图像生成质量,无需额外标注。
Retrieval-guided Cross-view Image Synthesis
- 通过对比学习构建跨视角语义嵌入空间,捕捉视图间相似性。
- 在CVUSA、CVACT和VIGOR-GEN上实现更高检索准确率(R@1)与生成质量(FID)。
- 新数据集VIGOR-GEN聚焦城市场景复杂视角变化,适合视觉合成研究者。
信息检索技术在通过强大特征表示识别跨领域语义相似性方面表现出色,但在引导生成任务(尤其是跨视角图像合成)方面的潜力尚未充分探索。跨视角图像合成面临不同视角间建立可靠对应关系的重大挑战。为此,我们提出一种新颖的检索引导框架,重新定义检索技术如何促进有效的跨视角图像合成。与依赖语义分割图或预处理模块的现有方法不同,我们的框架通过对比学习训练,建立平滑的嵌入空间,捕捉不同视角间的语义相似性。此外,引入一种新型融合机制,利用这些嵌入引导图像合成,同时学习并编码视图不变与视图特有特征。为进一步推动该领域发展,我们提出了VIGOR-GEN,一个专注于城市场景、包含真实世界复杂视角变化的新数据集。大量实验表明,我们的检索引导方法在CVUSA、CVACT和VIGOR-GEN数据集上显著优于现有方法,尤其在检索准确率(R@1)和生成质量(FID)方面表现突出。本工作连接了信息检索与生成任务,为解决复杂跨域生成挑战提供了新思路。
原文摘要 · Abstract (English)
Information retrieval techniques have demonstrated exceptional capabilities in identifying semantic similarities across diverse domains through robust feature representations. However, their potential in guiding synthesis tasks, particularly cross-view image synthesis, remains underexplored. Cross-view image synthesis presents significant challenges in establishing reliable correspondences between drastically different viewpoints. To address this, we propose a novel retrieval-guided framework that reimagines how retrieval techniques can facilitate effective cross-view image synthesis. Unlike existing methods that rely on auxiliary information, such as semantic segmentation maps or preprocessing modules, our retrieval-guided framework captures semantic similarities across different viewpoints, trained through contrastive learning to create a smooth embedding space. Furthermore, a novel fusion mechanism leverages these embeddings to guide image synthesis while learning and encoding both view-invariant and view-specific features. To further advance this area, we introduce VIGOR-GEN, a new urban-focused dataset with complex viewpoint variations in real-world scenarios. Extensive experiments demonstrate that our retrieval-guided approach significantly outperforms existing methods on the CVUSA, CVACT and VIGOR-GEN datasets, particularly in retrieval accuracy (R@1) and synthesis quality (FID). Our work bridges information retrieval and synthesis tasks, offering insights into how retrieval techniques can address complex cross-domain synthesis challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。