让AI自动决定何时搜图、在哪插入,提升图文回答准确性
VIG-RL: Learning to Search and Insert for Verified Image Grounding

- 用强化学习驱动搜索-选择-插入的动态决策流程
- 在多个数据集上超越现有静态方法,显著提升图文对齐效果
- 适合需要精准图文结合的应用,如智能客服与教育助手
在知识密集型场景中,提供可靠的交错文本-图像回应需要精确整合检索到的真实视觉证据,即验证图像定位(VIG)。现有检索增强框架主要依赖解耦的静态流程,无法动态判断何时需要外部知识以及如何在上下文中恰当地插入视觉内容。为此,我们提出 VIG-RL,一种自主代理框架,将搜索-选择-插入流程建模为一个主动决策过程。该框架在动态 ReAct 风格循环中运行,通过强化学习优化,并由复合奖励系统指导,全面评估代理每一步工具执行和最终多模态对齐效果。大量实验表明,VIG-RL 建立了新基准,显著优于现有的静态基线方法。
原文摘要 · Abstract (English)
In knowledge-intensive scenarios, providing reliable interleaved text-image responses requires Verified Image Grounding (VIG): the precise integration of retrieved authentic visual evidence. Existing retrieval-augmented frameworks predominantly rely on decoupled, static pipelines, inherently failing to dynamically reason about when external knowledge is required and where visual assets should be contextually inserted. To bridge this gap, we propose VIG-RL, an autonomous agentic framework that formulates the search-selection-insertion workflow as an active decision-making process. Operating within a dynamic ReAct-style loop, VIG-RL is optimized via reinforcement learning, guided by a composite reward system that holistically evaluates the agent's step-by-step tool execution and final multimodal alignment. Extensive evaluations demonstrate that VIG-RL establishes a new state-of-the-art, significantly outperforming existing static baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。