arXiv:2410.05601cs.CV2024-10NeurIPS被引 23

用检索外部图像解决图像修复模型的幻觉问题。

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

  • 通过检索高质量参考图,补充模型内部知识不足。
  • 无需训练即可提升修复细节真实性和保真度。
  • 适用于各类现有修复模型,通用性强。

基于扩散的大型修复模型(LRMs)通过模型权重中的内嵌知识显著提升了照片级图像修复效果。然而,现有LRMs在处理严重退化时常因过度依赖有限的内部知识而产生错误内容或纹理,陷入幻觉困境。本文提出一种正交解决方案——检索增强型图像修复框架(ReFIR),通过引入检索到的高质量图像作为外部知识,扩展现有LRMs的知识边界,生成更符合原始场景的细节。具体地,先采用最近邻查找检索内容相关的高质量参考图像,再设计跨图像注入机制,使LRMs能够利用这些参考图像中的优质纹理。得益于额外的外部知识,ReFIR有效缓解了幻觉问题,实现了高保真且真实的修复结果。大量实验表明,ReFIR无需训练,可适配多种现有LRMs,具有良好的通用性。

原文摘要 · Abstract (English)

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents or textures when dealing with severe degradations, due to their heavy reliance on limited internal knowledge. In this paper, we propose an orthogonal solution called the Retrieval-augmented Framework for Image Restoration (ReFIR), which incorporates retrieved images as external knowledge to extend the knowledge boundary of existing LRMs in generating details faithful to the original scene. Specifically, we first introduce the nearest neighbor lookup to retrieve content-relevant high-quality images as reference, after which we propose the cross-image injection to modify existing LRMs to utilize high-quality textures from retrieved images. Thanks to the additional external knowledge, our ReFIR can well handle the hallucination challenge and facilitate faithfully results. Extensive experiments demonstrate that ReFIR can achieve not only high-fidelity but also realistic restoration results. Importantly, our ReFIR requires no training and is adaptable to various LRMs.

图像修复检索增强扩散模型幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。