arXiv:2502.09411cs.CVcs.GR2025-02被引 23

用检索动态引导生成,让模型轻松画出罕见或细节复杂的图像。

ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation

  • 根据文本提示实时检索相关图像作为生成上下文。
  • 在不同基础模型上均显著提升罕见概念的生成质量。
  • 无需额外训练,可适配多种现有图像生成模型。

扩散模型能生成高质量且多样化的视觉内容,但在生成罕见或未见过的概念时表现不佳。为此,我们探索将检索增强生成(RAG)应用于图像生成模型。提出ImageRAG方法:根据给定文本提示动态检索相关图像,并将其作为上下文指导生成过程。与以往需专门训练检索增强生成模型的方法不同,ImageRAG直接利用现有图像条件模型的能力,无需进行RAG专用训练。该方法高度灵活,可适用于多种模型类型,在不同基线模型上均显著提升对罕见和细粒度概念的生成效果。

原文摘要 · Abstract (English)

Diffusion models enable high-quality and diverse visual content synthesis. However, they struggle to generate rare or unseen concepts. To address this challenge, we explore the usage of Retrieval-Augmented Generation (RAG) with image generation models. We propose ImageRAG, a method that dynamically retrieves relevant images based on a given text prompt, and uses them as context to guide the generation process. Prior approaches that used retrieved images to improve generation, trained models specifically for retrieval-based generation. In contrast, ImageRAG leverages the capabilities of existing image conditioning models, and does not require RAG-specific training. Our approach is highly adaptable and can be applied across different model types, showing significant improvement in generating rare and fine-grained concepts using different base models. Our project page is available at: https://rotem-shalev.github.io/ImageRAG

图像生成检索增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。