揭露关系抽取基准的不透明问题,呼吁更严谨的评估标准
Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards
- 分析主流基准数据集创建过程中的文档缺失问题
- 发现TACRED和NYT数据集存在严重类别不平衡与标签噪声
- 建议增加细粒度性能指标,提升模型评估透明度
本文研究自然语言处理中关系抽取(RE)任务的基准构建与排行榜透明度问题。现有RE基准普遍存在文档不完整,缺少数据来源、标注者一致性、实例筛选算法及潜在偏差(如数据集不平衡)等关键信息。当前进展多依赖排行榜,仅以F1-score等聚合指标排名,缺乏细粒度性能分析,难以反映模型的真实泛化能力。我们的分析显示,TACRED和NYT等常用基准存在显著类别不平衡与标签噪声,且缺乏基于类别的性能指标,无法准确评估多关系类型数据集上的模型表现。这些问题应在报告关系抽取进展时予以重视。尽管聚焦于关系抽取,但观察结果对其他NLP任务亦具普遍意义。本文主张改进文档记录与评估方法,而非否定现有基准与模型的价值,以推动领域健康发展。
原文摘要 · Abstract (English)
This paper investigates the transparency in the creation of benchmarks and the use of leaderboards for measuring progress in NLP, with a focus on the relation extraction (RE) task. Existing RE benchmarks often suffer from insufficient documentation, lacking crucial details such as data sources, inter-annotator agreement, the algorithms used for the selection of instances for datasets, and information on potential biases like dataset imbalance. Progress in RE is frequently measured by leaderboards that rank systems based on evaluation methods, typically limited to aggregate metrics like F1-score. However, the absence of detailed performance analysis beyond these metrics can obscure the true generalisation capabilities of models. Our analysis reveals that widely used RE benchmarks, such as TACRED and NYT, tend to be highly imbalanced and contain noisy labels. Moreover, the lack of class-based performance metrics fails to accurately reflect model performance across datasets with a large number of relation types. These limitations should be carefully considered when reporting progress in RE. While our discussion centers on the transparency of RE benchmarks and leaderboards, the observations we discuss are broadly applicable to other NLP tasks as well. Rather than undermining the significance and value of existing RE benchmarks and the development of new models, this paper advocates for improved documentation and more rigorous evaluation to advance the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。