用工具主动验证多模态推理,防止视觉幻觉和逻辑错误。
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM
- 引入工具辅助的主动验证机制,避免盲目认同错误答案。
- 8B模型在VisualProcessBench上超越72B、78B等更大模型。
- 支持可解释的验证过程,适合需要可靠推理的科研与应用。
多模态大语言模型在数学推理中表现优异,但仍易受视觉幻觉和逻辑不一致的影响,传统基于结果的监督无法有效缓解。虽然过程奖励模型(PRM)可实现逐步验证,但现有方法多为标量评分器或生成式批评者,存在谄媚问题,盲目认可错误假设而非基于视觉事实进行判断。为此,我们提出TIM-PRM(工具集成多模态PRM),一种新型智能体框架,将验证从被动分类转变为主动、工具增强的调查过程。TIM-PRM显式规划验证策略,并采用独立提问机制,通过外部工具查询证据,有效解耦验证与推理上下文,消除确认偏误。我们构建了高质量的工具集成验证轨迹数据集。在VisualProcessBench上的大量实验表明,我们的8B参数模型显著优于现有开源多模态PRM,性能超越更大的Qwen2.5-72B和InternVL-78B模型,同时提供可解释的验证过程洞察。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved impressive performances in mathematical reasoning, yet they remain vulnerable to visual hallucinations and logical inconsistencies that standard outcome-based supervision fails to mitigate. While Process Reward Models (PRMs) promise step-by-step verification, current approaches typically operate as scalar scorers or generative critics that suffer from sycophancy, blindly validating the flawed hypotheses rather than grounding them in visual reality. To bridge this gap, we introduce TIM-PRM (Tool-Integrated Multimodal PRM), a novel agentic framework that transforms verification from a passive classification task into an active, tool-augmented investigation. TIM-PRM is trained to explicitly plan verification strategies and utilizes a mechanism of Independent Question Asking to query evidence via external tools, effectively decoupling verification from the reasoning context to eliminate confirmation bias. We instantiate this method by curating a high-quality dataset of tool-integrated verification trajectories. Extensive experiments on VisualProcessBench demonstrate that our 8B parameter model surpasses existing open-source multimodal PRMs, significantly outperforming much larger models like Qwen2.5-72B and InternVL-78B, while offering interpretable insights into the verification process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。