arXiv:2503.01222cs.CVcs.CL2025-03ICML被引 26

用检索增强技术提升大模型对高分辨率图像的理解能力

Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

  • 不依赖训练,通过检索并融合图像局部区域来增强感知
  • 在V* Bench上使模型性能提升43%,HR-Bench上提升19%
  • 适合需要精准理解高清图像的视觉任务研究者

高分辨率图像感知仍是多模态大语言模型(MLLMs)的关键挑战。为克服现有方法局限,本文摒弃以往专用启发式策略,重新聚焦于提升MLLM长上下文处理能力,受通用大语言模型中检索增强生成(RAG)技术进展启发,首次探索将RAG应用于高分辨率感知问题。为此,提出无需训练的Retrieval-Augmented Perception(RAP)框架,通过空间感知布局保留空间上下文信息,动态检索并融合相关图像裁片。针对不同任务需求,提出的检索-探索搜索(RE-Search)机制根据模型置信度与检索得分自适应选择最优裁片数量。在高分辨率基准测试上的实验表明,RAP显著有效,使LLaVA-v1.5-13B在V* Bench上性能提升43%,在HR-Bench上提升19%。

原文摘要 · Abstract (English)

High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To overcome the limitations of existing methods, this paper shifts away from prior dedicated heuristic approaches and revisits the most fundamental idea to HR perception by enhancing the long-context capability of MLLMs, driven by recent advances in long-context techniques like retrieval-augmented generation (RAG) for general LLMs. Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench.

图像感知视觉RAG高分辨率多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。