让AI主动看图验证,提升多模态检索准确性
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
- 用可调用视觉工具的智能体,边推理边查图
- 平均检索准确率提升23.0%,在模糊图像上更可靠
- 适合需要精准视觉理解的跨模态任务
多模态大语言模型(MLLMs)被用于通用多模态检索,通过思维链(CoT)推理改进候选结果重排序。然而,现有方法仍以语言驱动为主,依赖静态视觉编码,缺乏主动验证细粒度视觉证据的能力,常在视觉模糊情况下产生推测性推理。我们提出V-Retrver,一种基于视觉证据的检索框架,将多模态检索重构为以视觉检查为基础的智能体推理过程。V-Retrver使MLLM能通过外部视觉工具在推理中选择性获取视觉证据,实现假设生成与目标化视觉验证之间的多模态交错推理。为训练该证据收集检索智能体,我们采用基于课程的学习策略,结合监督式推理激活、拒绝式优化及以证据对齐为目标的强化学习。在多个多模态检索基准上的实验表明,该方法在检索准确率上平均提升23.0%,显著增强感知驱动推理的可靠性与泛化能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain-of-Thought (CoT) reasoning improves candidate reranking. However, existing approaches remain largely language-driven, relying on static visual encodings and lacking the ability to actively verify fine-grained visual evidence, which often leads to speculative reasoning in visually ambiguous cases. We propose V-Retrver, an evidence-driven retrieval framework that reformulates multimodal retrieval as an agentic reasoning process grounded in visual inspection. V-Retrver enables an MLLM to selectively acquire visual evidence during reasoning via external visual tools, performing a multimodal interleaved reasoning process that alternates between hypothesis generation and targeted visual verification.To train such an evidence-gathering retrieval agent, we adopt a curriculum-based learning strategy combining supervised reasoning activation, rejection-based refinement, and reinforcement learning with an evidence-aligned objective. Experiments across multiple multimodal retrieval benchmarks demonstrate consistent improvements in retrieval accuracy (with 23.0% improvements on average), perception-driven reasoning reliability, and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。