arXiv:2605.20537cs.CL2026-05ACL

剖析医学命名实体识别基准数据集的真实差异,揭示其背后隐藏的评估偏差。

What Do Biomedical NER and Entity Linking Benchmarks Measure? A Corpus-Centric Diagnostic Framework

论文配图:What Do Biomedical NER and Entity Linking Benchmarks Measure? A Corpus-Centric Diagnostic Framework
图 1 · 摘自论文原文
  • 从语料标注、概念链接等角度构建诊断框架,量化分析数据集特性
  • 发现同类任务数据集在覆盖范围、训练测试重叠上差异显著
  • 适合关注医学文本模型泛化能力的研究者和评估者使用

生物医学命名实体识别(NER)和实体链接(EL)高度依赖标注语料,但这些资源在基准测试中的有效性常被假设而非验证。本文提出一种以语料为中心的诊断框架,直接基于语料标注、概念链接、训练测试划分、文档元数据及术语映射,系统分析九个涵盖疾病、化学物质和细胞类型的语料。框架将标准化统计归纳为五类:(1)规模、密度与标签分布,(2)词汇与概念结构,(3)训练测试重叠,(4)元数据构成,(5)术语覆盖率。分析显示,即使任务看似相同,不同语料在评估信号、泛化要求、训练测试复用程度及文献与概念空间覆盖上存在显著差异。这表明仅靠语料大小、实体类型等表面指标无法充分刻画基准评估内容。本文主张语料中心诊断可超越表层描述,识别潜在迁移风险,并解释基准结论的适用范围。框架代码及交互式仪表板已开源,支持复现分析与扩展语料评估。

原文摘要 · Abstract (English)

Biomedical named entity recognition (NER) and entity linking (EL) strongly depend on annotated corpora, but the utility of these resources for benchmarking is often assumed rather than characterized. We present a corpus-centric framework for diagnosing benchmark-relevant properties directly from corpus annotations, concept links, train-test splits, document metadata, and terminology mappings. The framework organizes standardized statistics into five families: (1) scale, density and label distribution, (2) lexical and conceptual structure, (3) train-test overlap, (4) metadata composition, and (5) terminology coverage where applicable. Applying the framework to nine corpora spanning diseases, chemicals, and cell types, we find that corpus properties can differ substantially, even when they address the same apparent task. We find differences in the evaluation signal they provide, the generalization demands they impose, the degree of train-test reuse they permit, and the regions of biomedical literature and concept space they represent. These differences suggest that commonly reported corpus statistics can be insufficient to characterize what biomedical NER and EL benchmarks evaluate. We argue that corpus-centric diagnostics provide a practical framework for analyzing corpora beyond surface descriptors such as corpus size and entity type, for identifying potential transfer risks, and for interpreting the scope of benchmarking conclusions. We release the framework as open-source code with an interactive dashboard to support reproducing our analyses and characterizing additional corpora.

医学NLP数据集评估实体识别语料诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。