arXiv:2503.06553cs.AIcs.CV2025-03ICCV被引 9

首个评估多模态模型推理过程的基准,助力提升科学问题判断可靠性。

ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges

  • 构建跨学科多模态评测集,每步标注正确性与错误类型。
  • 开源模型在过程判断上表现显著落后于闭源模型。
  • 提供17.3万条指令数据与双阶段微调策略,提升开源模型能力。

由于多模态大语言模型(MLLM)在解决科学问题时常出现错误,评估其推理过程的有效性对保障可靠性、发现细微缺陷至关重要。人工评估成本高,因此常采用让MLLM充当自动化过程裁判。然而,这类模型裁判的可靠性仍存疑。为此,我们提出ProJudgeBench,首个专为评估MLLM类过程裁判能力设计的综合性基准。该基准包含2,400个测试用例和50,118条步骤级标签,覆盖四个科学领域,具有多样难度与多模态内容。每个步骤均由专家标注正确性、错误类型及解释,支持系统性评估裁判模型在错误检测、分类与诊断方面的能力。在ProJudgeBench上的评估显示,开源模型与闭源模型间存在显著性能差距。为缩小这一差距,我们进一步提出ProJudge-173k大规模指令微调数据集,以及一种动态双阶段微调策略,促使模型先显式推理再评估解法。两项贡献显著提升了开源模型的过程评估能力。所有资源将公开,以推动可靠多模态过程评估研究。

原文摘要 · Abstract (English)

As multi-modal large language models (MLLMs) frequently exhibit errors when solving scientific problems, evaluating the validity of their reasoning processes is critical for ensuring reliability and uncovering fine-grained model weaknesses. Since human evaluation is laborious and costly, prompting MLLMs as automated process judges has become a common practice. However, the reliability of these model-based judges remains uncertain. To address this, we introduce ProJudgeBench, the first comprehensive benchmark specifically designed for evaluating abilities of MLLM-based process judges. ProJudgeBench comprises 2,400 test cases and 50,118 step-level labels, spanning four scientific disciplines with diverse difficulty levels and multi-modal content. In ProJudgeBench, each step is meticulously annotated by human experts for correctness, error type, and explanation, enabling a systematic evaluation of judges' capabilities to detect, classify and diagnose errors. Evaluation on ProJudgeBench reveals a significant performance gap between open-source and proprietary models. To bridge this gap, we further propose ProJudge-173k, a large-scale instruction-tuning dataset, and a Dynamic Dual-Phase fine-tuning strategy that encourages models to explicitly reason through problem-solving before assessing solutions. Both contributions significantly enhance the process evaluation capabilities of open-source models. All the resources will be released to foster future research of reliable multi-modal process evaluation.

多模态过程评估指令微调基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。