arXiv:2604.25923cs.CL2026-04

梳理NLP评估的常见问题,构建系统性分类框架。

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing

  • 提出NLP评估问题的分类体系,整合长期争议点。
  • 归纳评估方法中的核心权衡,如准确率与可解释性。
  • 提供检查清单,帮助设计更严谨的评估方案。

大语言模型的进展引发了对现有评估方法的广泛质疑。然而,这些讨论在自然语言处理领域已有长期积累。本文开展了一项范围综述,系统梳理了NLP评估研究中的核心关切,并构建了一个分类体系,整合各领域的重复观点与权衡。同时,讨论该分类的实际应用,提出一个结构化检查清单,以支持更审慎的评估设计与解读。通过将当前争论置于历史背景中,本工作为评估实践提供了综合性参考。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have prompted a growing body of work that questions the methodology of prevailing evaluation practices. However, many such critiques have already been extensively debated in natural language processing (NLP): a field with a long history of methodological reflection on evaluation. We conduct a scoping review of research on evaluation concerns in NLP and develop a taxonomy, synthesizing recurring positions and trade-offs within each area. We also discuss practical implications of the taxonomy, including a structured checklist to support more deliberate evaluation design and interpretation. By situating contemporary debates within their historical context, this work provides a consolidated reference for reasoning about evaluation practices.

评估方法NLP分类体系反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。