arXiv:2605.28375cs.CL2026-05ACL

首个针对朊病毒病的医学文献命名实体识别数据集,助力罕见病信息抽取。

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

论文配图:PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature
图 1 · 摘自论文原文
  • 人工标注317篇文献,覆盖15类粗粒度、31类细粒度临床实体
  • 标注一致性达81.78%精确匹配F1,结构复杂实体识别仍具挑战
  • 适合罕见病、精准医疗与低资源场景下的生物医学NLP研究

朊病毒病是罕见、进展迅速且致命的神经退行性疾病,早期诊断困难,因临床表现不特异。目前尚无公开的专注朊病毒病的实体标注数据集。本文提出PrionNER,一个在PubMed摘要中手动标注的命名实体识别数据集,包含317篇文献、2,943句、6,955个文本绑定实体,涵盖疾病、症状、诊断、发现、解剖、治疗及时间统计证据等15类粗粒度和31类细粒度临床相关实体类型。标注者间一致性达到81.78%精确匹配F1,表明标注一致性强。我们对BERT基线、W2NER及零样本提取器进行了基准测试,其中W2NER为最优监督模型,Gemma-4-31B为最强零样本模型,但对结构复杂提及和细粒度标签区分仍具挑战。PrionNER为朊病毒病信息抽取提供了临床基础基准,支持低资源、细粒度、非平铺式抽取条件下的罕见病生物医学NLP研究。数据集、标注指南与评估脚本已开源:https://github.com/daotuanan/PrionNER/

原文摘要 · Abstract (English)

Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stages because of nonspecific clinical presentations. However, to our knowledge, there is no publicly available prion-disease-focused dataset designed to capture a broad range of clinically relevant entities from the biomedical literature. We introduce PrionNER, a manually annotated named entity recognition dataset for prion disease clinical information in PubMed abstracts. The current release comprises 317 abstracts, 2,943 sentences, and 6,955 text-bound entity annotations spanning 15 coarse-grained and 31 fine-grained clinically oriented entity types covering diseases, symptoms, diagnostics, findings, anatomy, treatments, and temporal and statistical evidence. Inter-annotator agreement reaches 81.78 exact-match F1, indicating strong annotation consistency. We benchmark supervised BERT baselines, W2NER, and zero-shot extractors on PrionNER. W2NER is the strongest supervised model, and Gemma-4-31B is the strongest zero-shot model, but the benchmark remains challenging, especially for structurally complex mentions and fine-grained clinically adjacent label distinctions. PrionNER provides a clinically grounded benchmark for prion-disease information extraction and supports research on rare-disease biomedical NLP under low-resource, fine-grained, and non-flat extraction conditions. The dataset, annotation guidelines, and evaluation scripts are available at https://github.com/daotuanan/PrionNER/.

命名实体识别罕见病生物医学NLP数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。