arXiv:2508.04576cs.AI2025-08被引 1

首个评估多模态模型推理信心可靠性的基准,发现现有模型信心不可靠。

ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges

  • 构建三类对抗性干扰的推理步骤,测试模型信心鲁棒性
  • 14个主流模型信心表现差,存在严重校准偏差
  • 提供新指标与基线,适合模型可靠性研究者使用

推理是多模态大模型解决复杂任务的关键能力,判断推理步骤正确性对提升该能力至关重要。近年来,基于多模态大模型的推理过程判别器(MPJs)被广泛用于评估多模态任务中推理步骤的正确性。因此,评估MPJs的表现对于发现其局限性并指导改进具有重要意义。然而,现有基准主要关注步骤正确性分类和推理过程搜索,忽视了一个关键问题:MPJs在步骤层面生成的信心分数是否可靠。为填补这一空白,我们提出ConfProBench,首个系统评估MPJs步骤级信心可靠性的综合性基准。该基准构造了三类对抗性扰动的推理步骤:同义词替换、句法变换和图像扰动,以测试信心在扰动下的鲁棒性。同时,引入三项新指标:信心鲁棒性得分(CRS)、信心敏感性得分(CSS)和信心校准得分(CCS),分别评估鲁棒性、敏感性和校准性。我们评估了14个最先进的多模态大模型,包括商用和开源模型。实验揭示了当前MPJs在信心表现上的显著局限,并提供了竞争性基线以支持未来研究。

原文摘要 · Abstract (English)

Reasoning is a critical capability of multimodal large language models (MLLMs) for solving complex multimodal tasks, and judging the correctness of reasoning steps is crucial for improving this capability. Recently, MLLM-based process judges (MPJs) have been widely used to assess the correctness of reasoning steps in multimodal tasks. Therefore, evaluating MPJs is important for identifying their limitations and guiding future improvements. However, existing benchmarks for MPJs mainly focus on tasks such as step correctness classification and reasoning process search, while overlooking a key aspect: whether the confidence scores produced by MPJs at the step level are reliable. To address this gap, we propose ConfProBench, the first comprehensive benchmark designed to systematically evaluate the reliability of step-level confidence scores generated by MPJs. Our benchmark constructs three types of adversarially perturbed reasoning steps: Synonym Substitution, Syntactic Transformation, and Image Perturbation, to test the robustness of MPJ confidence under perturbations. In addition, we introduce three novel evaluation metrics: Confidence Robustness Score (CRS), Confidence Sensitivity Score (CSS), and Confidence Calibration Score (CCS), which evaluate robustness, sensitivity, and calibration, respectively. We evaluate 14 state-of-the-art MLLMs, including both proprietary and open-source models. Experiments reveal limitations in current MPJs' confidence performance and offer competitive baselines to support future research.

多模态模型评估信心评分基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。