提出评测多模态模型裁判偏见的基准,发现其评估易受模态忽略和不均衡影响。
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

- 构建可控扰动的多模态评测数据集,识别九类组合偏见
- 26个顶尖模型测试显示普遍模态忽视与评价失衡
- 适合关注AI评估可靠性的研究者与开发者使用
多模态大语言模型(MLLM)正被广泛用作自动评估器,即“MLLM-as-a-Judge”范式。然而其可靠性及对偏见的敏感性仍缺乏系统研究。我们发现,许多MLLM裁判无法稳定整合关键视觉或文本线索,在证据缺失或不匹配时产生不可靠评估,并对语义无关扰动表现出不稳定性。为此,我们系统定义了MLLM-as-a-Judge中的组合偏见,提出MM-JudgeBias基准。该基准在查询、图像和响应层面引入受控扰动,通过两个互补指标——偏差偏离(BD)衡量敏感性,偏差一致性(BC)衡量稳定性——评估模型行为。数据集包含超过1,800个精心筛选的多模态样本,源自29个原始基准,支持对九类偏见在多样化任务与领域中的细粒度诊断。对26个先进MLLM的实验揭示系统性模态忽视与不对称评价倾向,凸显提升裁判可靠性的迫切需求。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliability and vulnerabilities to biases remain underexplored. We find that many MLLM judges fail to reliably integrate key visual or textual cues, yielding unreliable evaluations when evidence is missing or mismatched, and exhibiting instability under semantically irrelevant perturbations. To address this, we systematically define Compositional Bias in MLLM-as-a-Judge systems and introduce MM-JudgeBias, a benchmark for evaluating it. MM-JudgeBias introduces controlled perturbations across Query, Image, and Response, and evaluates model behavior via two complementary metrics: Bias-Deviation (BD) for sensitivity and Bias-Conformity (BC) for stability. Our dataset of over 1,800 curated and refined multimodal samples, drawn from 29 source benchmarks, enables a fine-grained diagnosis of nine bias types across diverse tasks and domains. Experiments on 26 state-of-the-art MLLMs reveal systematic modality neglect and asymmetric evaluation tendencies, underscoring the need for more reliable judges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。