不训练模型,通过分析多模态判断分歧来验证推理步骤正确性
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

- 利用多个冻结验证器间的分歧信号,构建可解析的共识机制
- 在六个基准上提升基线模型性能最高达5.95%,媲美需大量标注的专用模型
- 无需微调,适用于多种任务,适合追求高效验证的开发者
多模态大语言模型生成的推理链常含细微错误导致答案错误。现有验证方法存在明显局限:要么需要昂贵的标注监督且跨任务表现不一,要么仅用简单聚合多个来源评分,忽略了关键洞察——当评分产生分歧时,这种分歧本身正是判断推理步骤是否有效的关键信息。我们将其形式化为一组异构冻结验证器之间的耦合评分问题,可解释为具有唯一闭式解均衡的协调博弈:一致表示步骤有效,分歧揭示不稳定性。为此,提出无需训练的通用步骤级验证方法VERDICT(基于分歧感知耦合阈值的验证)。据我们所知,VERDICT是首个使跨模态分歧结构显式且可操作的无训练验证器。通过闭式解计算共识分,实现分歧感知过滤与稳定性敏感排序。在六个基准上评估,该方法相较基线模型性能提升最高达+5.95%,表现媲美需大量监督的领域专用批评者,证明跨模态一致性可在无需任务特化适应的情况下提供稳健的验证信号。
原文摘要 · Abstract (English)
Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with inconsistent cross-task performance or aggregate scores from multiple sources by simple aggregations, missing a key insight: when these scores disagree, that disagreement itself carries important information about whether a reasoning step is truly valid or not. We formalise this as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium where agreement signals valid steps while disagreement reveals instability. Towards this end, we propose a training-free domain-agnostic step-wise verification approach we call VERDICT: VERification via Disagreement-Informed Coupled Thresholding. To our knowledge, VERDICT is the first training-free verifier that makes the structure of cross-modal disagreement explicit and actionable. It computes consensus scores through a closed-form solution, enabling both disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, \method consistently improves over the base model by up to +5.95%, and performs competitively with domain-specific critics that demand extensive supervision, demonstrating that cross-modal agreement provides robust verification signals without task-specific adaptation and Training-Free Verification
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。