分析2020-2025年NLG论文评估趋势,揭示评测方法三大痛点。
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
- 用多模型信息抽取分析1.4万篇论文,发现评测方法仍依赖旧式指标
- 92%论文未做人工验证,大模型评分在细粒度维度上失效
- 提出最小评估清单,指导正确选指标和部署大模型评测
随着自然语言生成(NLG)在现代NLP中占据主导地位,可扩展的评估仍是关键瓶颈。因此,大模型作为评判者(LaaJ)的采用迅速增加,在2025年已超过人工评估。这一转变促使我们对当前评估实践进行批判性分析。我们克服了传统关键词过滤和人工审查的局限,采用多大模型信息提取流水线,从2020–2025年四大主流NLP会议的14,171篇论文中获取结构化元数据。分析3,334篇筛选后的NLG论文后,识别出三大系统性挑战:(1) 指标惯性:尽管生成任务趋向开放,但传统词汇指标(如BLEU、ROUGE)仍为主要指标,通常与语义指标并列使用而非替代;(2) 指标-标准映射问题:论文层面共现数据显示,通用自动指标被当作质量的广泛代理,却未明确其评估的具体生成维度;(3) 验证缺口:尽管LaaJ增长迅速,但配套的人工验证极少(少于8%)。关键的是,虽然LaaJ与整体质量相关,但在细粒度标准(如流畅性)上一致性崩溃。为此,我们基于发现提炼出一份最小评估检查清单,以指导指标选择、构建效度和LaaJ部署。
原文摘要 · Abstract (English)
As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly, appearing in more papers than human evaluation in 2025. This pivotal shift motivates a critical analysis of current evaluation practices. Overcoming the limits of rigid keyword filtering and manual review, we employ a multi-LLM information extraction pipeline to gather structured metadata from 14,171 papers across four major NLP conferences (2020-2025). Analyzing 3,334 filtered NLG papers, we identify three systemic challenges. (1) Metric inertia: despite the shift toward open-ended generation, legacy lexical metrics (BLEU, ROUGE) persist as primary indicators, typically used alongside rather than replaced by semantic alternatives. (2) Metric-criteria mapping problem: our paper-level co-occurrence data reveals that general-purpose automatic metrics are applied as broad proxies for quality, without specifying which dimension of text generation they are intended to evaluate. (3) Validation gap: LaaJ has grown rapidly without commensurate human validation (fewer than 8% of papers). Crucially, while LaaJ correlates with aggregate quality, alignment collapses on fine-grained criteria like fluency. To address these gaps, we distill our findings into a minimal Evaluation Checklist to guide metric selection, construct validity, and LaaJ deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。