arXiv:2602.04290cs.CL2026-02

让模型推理时有伙伴实时纠错,提升多模态问答准确率

Guided Verifier: Collaborative Multimodal Reasoning via Dynamic Process Supervision

  • 引入动态验证器与主模型协同解题,实时发现逻辑错误
  • 在MathVista等数据集上,8B模型性能接近更大规模模型
  • 专为幻觉问题设计数据集,支持推理过程的精准引导

强化学习已成为提升多模态大语言模型复杂推理能力的关键机制。然而,现有方法普遍采用孤立推演策略,模型独自完成推理,缺乏中间环节监督,导致早期逻辑偏差不断累积,引发不可逆失败,产生噪声优化信号。本文提出「Guided Verifier」框架,突破被动终端奖励模式,引入动态验证器与策略模型实时协同解题。在推演过程中,验证器与模型交互,检测不一致并提供方向性修正信号,引导其走向有效推理路径。为此,我们构建了针对多模态幻觉的专用数据合成管道,创建了包含过程级负样本的CoRe数据集和正确推理轨迹数据集CoRe,用于训练引导验证器。在MathVista、MathVerse和MMMU上的大量实验表明,通过分配计算资源进行协同推理与动态验证,一个8B参数模型可达到优异性能。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has emerged as a pivotal mechanism for enhancing the complex reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevailing paradigms typically rely on solitary rollout strategies where the model works alone. This lack of intermediate oversight renders the reasoning process susceptible to error propagation, where early logical deviations cascade into irreversible failures, resulting in noisy optimization signals. In this paper, we propose the \textbf{Guided Verifier} framework to address these structural limitations. Moving beyond passive terminal rewards, we introduce a dynamic verifier that actively co-solves tasks alongside the policy. During the rollout phase, this verifier interacts with the policy model in real-time, detecting inconsistencies and providing directional signals to steer the model toward valid trajectories. To facilitate this, we develop a specialized data synthesis pipeline targeting multimodal hallucinations, constructing \textbf{CoRe} dataset of process-level negatives and \textbf{Co}rrect-guide \textbf{Re}asoning trajectories to train the guided verifier. Extensive experiments on MathVista, MathVerse and MMMU indicate that by allocating compute to collaborative inference and dynamic verification, an 8B-parameter model can achieve strong performance.

多模态推理强化学习协同验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。