arXiv:2512.06373cs.CV2025-12被引 7

提出首个工具反馈优化的指代定位推理框架,解决视觉工具出错导致的幻觉问题。

VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning

  • 设计两阶段思考-重思机制,主动应对工具输出错误
  • 引入修正奖励机制,显著提升对错误工具结果的纠错能力
  • 适合需要高可靠性的多模态推理任务研究者

工具集成视觉推理(TiVR)在增强多模态问题求解方面展现出巨大潜力。然而,现有方法主要通过强化学习整合多种视觉工具,忽视了对不可靠或错误工具输出的有效应答机制。这一局限在指代与定位任务中尤为明显,错误的检测工具预测常导致模型生成幻觉推理。为此,我们提出首个面向工具反馈优化的指代定位推理框架——VG-Refiner。技术上,引入两阶段思考-重思机制,使模型能显式分析并响应工具反馈,并设计修正奖励以鼓励对劣质工具结果的有效修正。此外,提出两个新指标并建立公平评估协议,系统性衡量模型的修正能力。通过少量特定任务数据微调,VG-Refiner在指代与推理定位基准上显著提升准确率与纠错能力,同时保持预训练模型的通用性。

原文摘要 · Abstract (English)

Tool-integrated visual reasoning (TiVR) has demonstrated great potential in enhancing multimodal problem-solving. However, existing TiVR paradigms mainly focus on integrating various visual tools through reinforcement learning, while neglecting to design effective response mechanisms for handling unreliable or erroneous tool outputs. This limitation is particularly pronounced in referring and grounding tasks, where inaccurate detection tool predictions often mislead TiVR models into generating hallucinated reasoning. To address this issue, we propose the VG-Refiner, the first framework aiming at the tool-refined referring grounded reasoning. Technically, we introduce a two-stage think-rethink mechanism that enables the model to explicitly analyze and respond to tool feedback, along with a refinement reward that encourages effective correction in response to poor tool results. In addition, we propose two new metrics and establish fair evaluation protocols to systematically measure the refinement ability of current models. We adopt a small amount of task-specific data to enhance the refinement capability of VG-Refiner, achieving a significant improvement in accuracy and correction ability on referring and reasoning grounding benchmarks while preserving the general capabilities of the pretrained model.

视觉推理工具集成强化学习指代定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。