arXiv:2512.05111cs.CV2025-12被引 10

让奖励模型像人一样用工具查证,提升多模态推理的可信度。

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

  • 引入工具调用机制,自主验证视觉细节与文档信息
  • 在多任务基准上平均提升16.2%,工具使用任务提升9.6%
  • 适合需要高可信度推理的复杂多模态应用

奖励模型对对齐视觉-语言系统与人类偏好至关重要,但现有方法存在幻觉、视觉定位弱及无法调用工具验证的问题,限制了其在复杂多模态推理任务中的可靠性。我们提出ARM-Thinker,一种具备代理能力的多模态奖励模型,可自主调用外部工具(如图像裁剪、文档页检索)以基于可验证证据进行判断,取代静态、非交互式评分。该机制使模型能验证细粒度视觉特征、跨多页文档交叉比对,并验证推理结论,这些能力在现有模型中缺失。我们采用多阶段强化学习联合优化工具调用决策与判断准确性。为评估代理式奖励建模,我们构建ARMBench-VL,包含三个基准:细粒度视觉定位(图像级工具)、多页文档理解(检索工具)和指令遵循(文本级验证)。ARM-Thinker在奖励建模基准上平均提升16.2%,工具使用任务提升9.6%,并在多模态数学与逻辑推理基准上优于基线。结果表明,代理能力显著提升奖励模型的准确性和可解释性。

原文摘要 · Abstract (English)

Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to use tools for verification, limiting their reliability on complex multimodal reasoning tasks. We present ARM-Thinker, an A}gentic multimodal Reward Model that autonomously invokes external tools (e.g., image cropping, doc page retrieval) to ground judgments in verifiable evidence, replacing static, non-interactive reward scoring. This enables the model to verify fine-grained visual details, cross-reference multi-page evidence, and validate reasoning claims, which are capabilities absent in existing reward models. We train ARM-Thinker with multi-stage reinforcement learning, jointly optimizing tool-calling decisions and judgment accuracy. To evaluate agentic reward modeling, we introduce ARMBench-VL, comprising three benchmarks that assess fine-grained visual grounding (image-level tools), multi-page document understanding (retrieval tools), and instruction following (text-level verification). ARM-Thinker achieves +16.2% average improvement on reward modeling benchmarks, +9.6% on tool-use tasks, and outperforms baselines on multimodal math and logical reasoning benchmarks. Our results demonstrate that agentic capabilities significantly enhance both accuracy and interpretability of reward models.

多模态奖励模型代理智能视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。