构建多维评测基准与数据生成框架,提升视觉语言模型判别能力
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
- 设计十维能力评测基准,细分推理、长度偏差等子任务
- 通过强化学习生成多样推理路径,训练出更强判别模型
- 适用于评估和优化多模态判别模型的可靠性与一致性
使用多模态大语言模型(MLLM)作为评判者,在多个领域逐渐成为新范式。评估 MLLM 作为评判者的能力与可靠性对保障可信评估至关重要。现有评判基准按任务类型分类,但未能捕捉可靠评估所需的核心判断能力。本文提出 M-JudgeBench,一个十维能力导向的评测基准,将评估分解为成对思维链比较、长度偏差规避和过程错误检测等任务,涵盖十个细粒度子任务。该设计可诊断模型在不同推理风格、响应长度及跨模型差异下的可靠性。系统性评估揭示了现有 MLLM 作为评判者系统的系统性弱点。为此,我们进一步提出 Judge-MCTS,一种生成多样化正确性与长度推理轨迹的数据构造框架。基于此框架,我们构建了 MCTS 增强数据集,并训练出 M-Judger 系列强判别模型。大量实验表明,M-Judger 在现有评测基准及 M-JudgeBench 上均表现更优。整体工作通过 M-JudgeBench 与 Judge-MCTS 框架,为多模态评判模型评估建立更严谨的基础,推动未来评判模型评估与能力驱动训练研究。
原文摘要 · Abstract (English)
Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm across various domains. Evaluating the capability and reliability of MLLM-as-a-judge systems is therefore essential for ensuring trustworthy assessment. Existing judge benchmarks categorize samples by task types but fail to capture the fundamental judgment capabilities required for reliable evaluation. In this work, we introduce M-JudgeBench, a ten-dimensional capability-oriented benchmark designed to comprehensively assess the judgment abilities of MLLMs. Our benchmark decomposes evaluation into pairwise Chain-of-Thought (CoT) comparison, length bias avoidance, and process error detection tasks, jointly covering ten fine-grained subtasks. This design enables diagnosis of model reliability across reasoning styles, response lengths, and cross-model variations. Systematic evaluation uncovers the systematic weaknesses in existing MLLM-as-a-judge systems. To address this issue, we further propose Judge-MCTS, a data construction framework generating pairwise reasoning trajectories with various correctness and length. Using Judge-MCTS, we construct an MCTS-augmented dataset and train M-Judger, a series of strong judge models. Extensive experiments demonstrate the superiority of M-Judger on existing judge benchmarks as well as M-JudgeBench. Overall, our work establishes a more principled foundation for evaluating MLLM-as-a-judge through M-JudgeBench and Judge-MCTS framework, paving the way for future research on judge model evaluation and capability-driven judge training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。