arXiv:2607.27959cs.CVcs.IR2026-07被引 11

通过细粒度上下文学习,提升大模型在复杂图像检索中的表现。

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

论文配图:FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
图 1 · 摘自论文原文
  • 构建细粒度图文数据集,支持精准上下文建模。
  • 分两阶段微调:先理解上下文,再对齐查询与目标图像。
  • 在五项复杂任务中零样本表现领先,适合高精度图像检索场景。

由于多模态大语言模型(MLLMs)具备强大的泛化多模态处理与推理能力,其在通用图像检索任务中展现出巨大潜力。然而,现有研究忽视了细粒度上下文建模和解耦微调目标对复杂检索任务(如长文本到图像、视觉对话检索、组合图像检索,CIR)的提升作用。为此,本文提出自动化细粒度多模态五元组数据集构建流程,以及一种两阶段细粒度多模态微调策略。该数据集生成包含细粒度图像描述和修改文本的完整CIR数据集,支持细粒度上下文建模。新方法将微调过程解耦为两个阶段:(1) 细粒度上下文推理导向微调;(2) 细粒度检索导向微调,依次增强模型的上下文理解与查询-目标对齐能力。在涵盖五种不同复杂图像检索任务的数据集上进行的大量实验表明,该方法在零样本设置下显著优于现有方法,且采用更轻量级的MLLM骨干网络。

原文摘要 · Abstract (English)

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model's context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.

多模态图像检索细粒度大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。