梳理NLU诊断基准现状,呼吁建立统一评估标准
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- 系统整理英、阿、多语种NLU基准中的诊断数据集
- 发现缺乏统一的术语命名与语言现象覆盖标准
- 适合关注模型细粒度分析与评估体系构建的研究者
自然语言理解(NLU)是自然语言处理(NLP)的基础任务,其评估已成为近年研究热点,催生了大量基准测试。这些基准涵盖多种任务与数据集,通过公开排行榜评估预训练模型性能。尤其值得关注的是,部分基准包含诊断数据集,用于探究和细粒度分析广泛的语言现象。本文全面综述了现有的英语、阿拉伯语及多语言NLU基准,重点关注其诊断数据集所覆盖的语言现象。通过详细比较与分析,揭示了现有基准在评估能力上的优势与局限,并开展深入的错误分析。研究发现,当前缺乏宏观与微观分类的命名规范,也无标准应覆盖的语言现象集合。因此提出核心问题:为何不为NLU诊断评估基准建立类似工业界ISO的标准?通过对语言现象覆盖范围的深度对比,旨在为未来构建全球性语言现象层级体系提供支持。我们认为,建立诊断评估的标准化指标,将有助于更深入地比较不同模型在各类诊断基准上的表现。
原文摘要 · Abstract (English)
Natural Language Understanding (NLU) is a basic task in Natural Language Processing (NLP). The evaluation of NLU capabilities has become a trending research topic that attracts researchers in the last few years, resulting in the development of numerous benchmarks. These benchmarks include various tasks and datasets in order to evaluate the results of pretrained models via public leaderboards. Notably, several benchmarks contain diagnostics datasets designed for investigation and fine-grained error analysis across a wide range of linguistic phenomena. This survey provides a comprehensive review of available English, Arabic, and Multilingual NLU benchmarks, with a particular emphasis on their diagnostics datasets and the linguistic phenomena they covered. We present a detailed comparison and analysis of these benchmarks, highlighting their strengths and limitations in evaluating NLU tasks and providing in-depth error analysis. When highlighting the gaps in the state-of-the-art, we noted that there is no naming convention for macro and micro categories or even a standard set of linguistic phenomena that should be covered. Consequently, we formulated a research question regarding the evaluation metrics of the evaluation diagnostics benchmarks: "Why do not we have an evaluation standard for the NLU evaluation diagnostics benchmarks?" similar to ISO standard in industry. We conducted a deep analysis and comparisons of the covered linguistic phenomena in order to support experts in building a global hierarchy for linguistic phenomena in future. We think that having evaluation metrics for diagnostics evaluation could be valuable to gain more insights when comparing the results of the studied models on different diagnostics benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。