为自然语言评估标准建立统一命名与定义体系,解决结果不可比问题。
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
- 基于三项NLP评估调查,构建标准化评估准则层级体系。
- 明确不同实验中相同名称评估项的实际差异,提升可比性。
- 适用于评估对比、新实验设计及合规性审查,推动领域科学进步。
先前研究显示,相同质量评估名称(如流畅性)在不同实验中可能衡量不同方面,导致看似可比的结果实则不可比,阻碍了对系统质量的可靠判断,影响领域整体科学发展。为解决此问题,本文通过三轮评估调查,采用描述性方法提炼出一套标准评估准则名称与定义,并构建层级结构,使数百种现有名称得以映射与归一。该体系称为QCET,可用于:(i) 确立现有评估的可比性;(ii) 指导新评估设计;(iii) 评估监管合规性。该工作为实现跨研究、跨任务的客观评价提供基础支撑。
原文摘要 · Abstract (English)
Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quality, and the comparability implied by the name can be misleading. Not knowing when two evaluations are comparable in this sense means we currently lack the ability to draw reliable conclusions about system quality on the basis of multiple, independently conducted evaluations. This in turn hampers the ability of the field to progress scientifically as a whole, a pervasive issue in NLP since its beginning (Sparck Jones, 1981). It is hard to see how the issue of unclear comparability can be fully addressed other than by the creation of a standard set of quality criterion names and definitions that the several hundred quality criterion names actually in use in the field can be mapped to, and grounded in. Taking a strictly descriptive approach, the QCET Quality Criteria for Evaluation Taxonomy derives a standard set of quality criterion names and definitions from three surveys of evaluations reported in NLP, and structures them into a hierarchy where each parent node captures common aspects of its child nodes. We present QCET and the resources it consists of, and discuss its three main uses in (i) establishing comparability of existing evaluations, (ii) guiding the design of new evaluations, and (iii) assessing regulatory compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。