arXiv:2607.02235cs.CLcs.AI2026-07

LLM评鉴在多语言和低资源场景下存在可靠性问题,需谨慎使用。

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

  • 分析650篇论文,仅33篇关注多语言与低资源场景下的LLM评鉴
  • 发现评估结果不一致,且普遍过度依赖单一模型判断
  • 建议提升多语言评估的多样性与人类验证,避免盲目信任

LLM-as-a-Judge已成为自然语言生成任务主流评估范式,因其相比传统指标更接近人类判断,但目前主要集中在英语。近年来开始尝试扩展至多语言及低资源语言场景,然而大模型在低资源语言上表现有限,且常缺乏充分的人类验证。我们分析了ACL Anthology中涉及多语言与低资源语言的论文,发现650篇提及LLM-as-a-Judge的论文中仅有33篇聚焦此类场景。深入分析显示,评估结果不一致,研究者倾向于过度信任模型判断,且多数研究仅依赖单一裁判模型。为此,我们提出若干建议,以指导NLP社区更审慎地在多语言和低资源情境中应用LLM-as-a-Judge。

原文摘要 · Abstract (English)

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English. There are now attempts to extend LLM-as-a-Judge to multilingual settings including low-resource languages. However, LLMs have limited proficiency in low-resource languages, and there is often no adequate human validation in these settings. To highlight the scope of the problem and current practices, we explore the use of LLM-as-a-Judge evaluators in ACL Anthology papers focusing on multilingual settings and low-resource languages across a diverse set of tasks. Out of 650 papers mentioning LLM-as-a-judge, only 33 of them focus on low-resource or multilingual settings. Our in-depth analysis of these papers indicates inconsistent evaluation outcomes, a tendency to overtrust LLM judgments in multilingual settings, and the widespread reliance on a single judge model per study. To help the NLP community further, we conclude with recommendations about how to use LLM-as-a-Judge in multilingual and low-resource settings.

大模型评估多语言低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。