MLLM做评分存在模型偏好偏差,自评倾向明显。
MLLM-as-a-Judge Exhibits Model Preference Bias

- 通过解耦评分偏好与生成质量差异,量化模型间评分偏见。
- 12个MLLM共129万条图文描述数据中发现自评倾向和家族互评现象。
- 提出Pomms集成方法,有效降低偏见且保持性能,适合基准评测者使用。
利用多模态大语言模型(MLLM)进行自动评估,即“MLLM-as-a-Judge”,已被广泛用于衡量模型性能。若此类方法存在偏差,可能扭曲模型比较并影响以基准为导向的科研进展。然而,目前尚不清楚MLLM-as-a-Judge是否对特定MLLM生成的文本存在偏好或歧视。本研究提出Philautia-Eval,用于探究这种模型特异性偏好偏差。Philautia-Eval通过解耦评分偏好与生成质量差异,量化偏差程度。基于从12个MLLM收集的129万组图文描述-评分对,我们发现代表性MLLM普遍存在自评偏好。实验结果还表明,特定模型家族间存在相互偏好,可能源于共享的连接器和重叠的指令微调资源。最后,我们引入一种简单的MLLM集成方法Pomms。结果表明,Pomms能有效缓解模型特异性偏好偏差,同时保持性能。
原文摘要 · Abstract (English)
Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model performance. If such MLLM-as-a-Judge methods were biased, they could distort model comparisons and benchmark-driven scientific progress. However, it remains unclear to what extent MLLM-as-a-Judge methods favor or disfavor text generated by specific MLLMs. In this study, we propose Philautia-Eval to investigate such model-specific preference bias. Philautia-Eval quantifies the degree of the bias by disentangling preference tendencies from differences in generation quality. Using 1.29M caption-score pairs collected from 12 MLLMs, we found that representative MLLMs tend to exhibit self-preference bias. Moreover, experimental results indicate mutual preference bias within particular model families, which is potentially driven by reused connectors and overlapping instruction-tuning resources. Finally, we introduce a simple ensemble of MLLMs, Pomms. Our results demonstrated that Pomms effectively mitigated the model-specific preference bias while maintaining performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。