首个多语言评估基准,检验大模型评分器对非英语内容的公正性。
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
- 构建覆盖18语言的核心集与122语言的一致性集,专为多语言评估设计。
- 发现现有评分模型在非英语上表现差,低资源语言更易被误判。
- 适用于测试多语言大模型评估器的公平性与一致性,适合评测研究者使用。
随着大语言模型在非英语语种上生成流利内容的能力提升,精准评估这些输出变得愈发重要。然而,现有评估方法多依赖擅长英文评估的模型,未充分验证其对非英语文本的有效性。现有元评估基准也以英语为中心。为填补这一空白,我们提出MM-Eval,一个涵盖18种语言的核心子集和覆盖122种语言的语言一致性子集的多语言元评估基准。该基准并非简单翻译英文基准,而是针对多语言特有挑战设计。不同于仅关注成对比较排序准确率的基准,MM-Eval还评估跨语言的绝对分数一致性和公平性。结果表明,现有在英语中表现优异的评估模型在非英语上仍有显著提升空间,且对低资源语言存在不公平与不一致现象。我们通过与Best-of-N排名的相关性验证了MM-Eval的有效性,其相关性显著高于其他基准。数据集与代码已公开。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are now capable of producing fluent and coherent content in languages other than English, it is not imperative to precisely evaluate these non-English outputs. However, when assessing the outputs from mutlilingual LLMs, prior works often employed LLM based evaluators that excel at assessing English outputs, without a thorough examination of whether these evaluators could effectively assess non-English text as well. Moreover, existing benchmarks to test evaluator LLMs (referred to as "meta-evaluation benchmarks") are mostly English-centric. To bridge this gap and examine whether evaluator LLMs can reliably assess the outputs of multilingual LLMs, we introduce MM-Eval, a multilingual meta-evaluation benchmark comprising five core subsets covering 18 languages and a Language Consistency subset spanning 122 languages. A core attribute of MM-Eval is that, instead of merely translating existing English meta-evaluation benchmarks, it is designed with multilingual-specific challenges in mind. Additionally, unlike existing meta-evaluation benchmarks that focus solely on ranking accuracy over pairwise data, MM-Eval also evaluates the consistency and fairness of absolute score values across a wide range of languages. Our results show that existing evaluator LLMs that excel in English contexts have considerable room for improvement when assessing non-English outputs. Furthermore, we find that evaluators are unfair and inconsistent when evaluating lower-resourced languages. Finally, we validate MM-Eval by measuring its correlation with Best-of-N rankings, finding a significantly stronger correlation compared to other meta-evaluation benchmarks. We publicly release our benchmark and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。