arXiv:2508.13816cs.CLcs.AI2025-08中稿 · RANLP 2025被引 1

评估AI生成文本的指标一直不靠谱,没有万能标准。

The illusion of a perfect metric: Why evaluating AI's words is harder than it looks

  • 分析现有自动评估指标的原理与局限性
  • 发现各类指标在不同任务中表现差异大,且与人评相关性不稳定
  • 建议按任务选指标,并加强验证方法,别再迷信‘完美指标’

自然语言生成(NLG)的评估对AI落地至关重要,但长期面临挑战。尽管人工评估被视为金标准,却成本高、难扩展。为应对这一问题,研究者开发了多种自动评估指标(AEM),通过比较模型输出与人类参考文本,生成近似人类判断的分数。从早期的词法比对,到语义相似度模型,再到近期基于大模型的评价器,指标不断演进。然而,至今未出现公认的最优解,导致研究中使用标准不一。本文系统考察现有指标的方法论、优缺点、验证方式及其与人类判断的相关性,揭示出若干关键问题:指标仅捕捉文本质量的部分维度,效果随任务和数据集变化,验证流程缺乏规范,且与人评相关性不一致。尤为关键的是,这些挑战在最新的‘大模型作为评判者’(LLM-as-a-Judge)及检索增强生成(RAG)评估中依然存在。研究结果质疑‘完美指标’的可行性,主张根据任务需求选择指标,采用互补评估,并推动新指标强化验证方法。

原文摘要 · Abstract (English)

Evaluating Natural Language Generation (NLG) is crucial for the practical adoption of AI, but has been a longstanding research challenge. While human evaluation is considered the de-facto standard, it is expensive and lacks scalability. Practical applications have driven the development of various automatic evaluation metrics (AEM), designed to compare the model output with human-written references, generating a score which approximates human judgment. Over time, AEMs have evolved from simple lexical comparisons, to semantic similarity models and, more recently, to LLM-based evaluators. However, it seems that no single metric has emerged as a definitive solution, resulting in studies using different ones without fully considering the implications. This paper aims to show this by conducting a thorough examination of the methodologies of existing metrics, their documented strengths and limitations, validation methods, and correlations with human judgment. We identify several key challenges: metrics often capture only specific aspects of text quality, their effectiveness varies by task and dataset, validation practices remain unstructured, and correlations with human judgment are inconsistent. Importantly, we find that these challenges persist in the most recent type of metric, LLM-as-a-Judge, as well as in the evaluation of Retrieval Augmented Generation (RAG), an increasingly relevant task in academia and industry. Our findings challenge the quest for the 'perfect metric'. We propose selecting metrics based on task-specific needs and leveraging complementary evaluations and advocate that new metrics should focus on enhanced validation methodologies.

文本评估大模型评测指标可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。