arXiv:2508.18076cs.CL2025-08NeurIPS被引 41

质疑大模型当裁判的可靠性,提醒评估需更严谨

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

  • 用社会科学研究中的测量理论分析大模型评分的四大假设
  • 发现大模型裁判在摘要、标注、安全对齐中存在不可靠性
  • 适合关注AI评估方法论的研究者与从业者参考

自然语言生成(NLG)系统的评估仍是自然语言处理的核心挑战,尤其在大语言模型(LLM)作为通用工具兴起的背景下更为复杂。近期,大语言模型作为裁判(LLJ)被视为传统指标的有前景替代方案,但其有效性尚未经过充分检验。本文认为当前对LLJ的热衷可能过于仓促,因其应用已超过对其可靠性和有效性的严格审视。基于社会科学中的测量理论,我们识别并批判性评估了使用LLJ的四个核心假设:能否作为人类判断的代理、评估能力、可扩展性以及成本效益。我们分析这些假设如何因LLM、LLJ本身的局限性或当前NLG评估实践而受到挑战。为支持分析,我们考察了三个应用场景:文本摘要、数据标注和安全对齐。最后强调,必须建立更负责任的评估实践,以确保LLJ在领域内的日益作用推动而非阻碍NLG的发展。

原文摘要 · Abstract (English)

Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recently, large language models as judges (LLJs) have emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.

大模型评估测量理论LLM裁判NLG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。