arXiv:2606.13051cs.AI2026-06ACL

构建自免领域标注语料,提升疾病与抗体信息抽取准确率。

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

论文配图:AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction
图 1 · 摘自论文原文
  • 人工标注115篇文献,涵盖自免病、自身抗体等五类实体及关系
  • 微调后命名实体识别性能显著提升,验证小规模标注有效性
  • 适合生物医学信息抽取研究者使用,尤其关注自免领域

尽管深度学习和大语言模型推动了信息抽取发展,但在高度专业的生物医学领域仍存在性能差距,因领域特异性复杂性对通用模型构成挑战。本文聚焦自免疫领域,关注自免疫疾病、自身抗体(可能标记或引发疾病的分子)、其分子靶点、体内位置及关联临床症状。我们提出AAbAAC(AutoAntibodies and Autoimmunity Annotated Corpus),一个从PubMed选取的115篇摘要构成的语料库,包含人工标注的实体及其关系。首先用于评估多种命名实体识别(NER)方法,其次用于微调NER模型。研究表明,使用AAbAAC进行微调后,NER性能明显提升,证明小规模标注在专业领域的价值,并推动自免疫的计算研究。该语料库已公开于https://github.com/f-maury/AAbAAC。

原文摘要 · Abstract (English)

Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models. In this work, we focus on the domain of autoimmunity, where the main entities of interest are autoimmune diseases, autoantibodies (i.e., molecules that may mark or cause these diseases), their molecular targets, their location in the body, and their associated clinical signs. Herein, we present AAbAAC (AutoAntibodies and Autoimmunity Annotated Corpus), a corpus of 115 abstracts selected from PubMed, where we manually annotated entities and their relationships. First, AAbAAC was used to evaluate several methods on the task of named entity recognition (NER), and secondly, to fine-tune NER models. Our study demonstrates the utility of AAbAAC for information extraction in the domain of autoimmunity, showing expected improvement in NER performance after finetuning. This illustrates the value of small-scale annotation efforts for specialized domains and contributes to the computational study of autoimmunity. The AAbAAC corpus is available at https://github.com/f-maury/AAbAAC.

信息抽取自免疾病标注语料NER

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。