arXiv:2607.14480cs.CL2026-07被引 1

大模型评估器对低资源语言存在系统性偏见,评分更宽松却难被发现。

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

论文配图:Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
图 1 · 摘自论文原文
  • 用跨23种语言的相同指令-回复对测试,发现不同语言得分差异显著。
  • 低资源语言内容通过安全过滤的概率比高资源语言高出43%。
  • 即使模型不自信时也更倾向给低资源语言高分,是结构性语言偏差。

LLM评估器(训练的奖励模型和提示式大模型作为裁判)通常通过成对准确率验证。在多语言场景下,这一方法基于高成对准确率即代表可靠、无语言偏见评分的假设。我们证明该假设不成立:在23种语言中使用语义相同的指令-响应对进行实验,发现多语言评估器对不同语言的评分存在显著差异。这种偏差在八种不同架构与训练方式的开源评估器中均显著且一致,甚至在前沿裁判中仍存在,并与语言资源水平强相关——低资源语言得分更优。然而,这种偏差在成对准确率中不可见:评估器成对准确率超过90%,但采用全局阈值时,不同语言的接受率差距高达43%。这意味着低资源语言中的有害内容更可能通过安全过滤。按语言设定阈值虽可缓解,但易被混码提示规避。进一步分析发现,模型不确定性与高分倾向相关:模型在不确定时更倾向于给出高分,无论使用负对数似然还是无标记不确定性度量;但控制不确定性后,语言身份仍是显著预测因子,偏差不能仅由内容难度解释,而是结构性的语言层面错配。

原文摘要 · Abstract (English)

LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.

大模型评估语言偏见安全过滤低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。