专用于免疫疾病文本的实体识别模型,提升医学文献解析精度。
Specialty-Specific Medical Language Model for Immune-Mediated Diseases

- 基于临床嵌入与专家标注构建领域专用模型
- 在12类实体上达到0.89的F1分数,优于通用模型
- 适合免疫与感染病研究者做病历分析与数据挖掘
从自由文本医学叙述中提取详细临床信息仍是研究者和医疗系统面临的实际挑战。免疫相关与感染性疾病术语在不同来源间差异显著,常导致通用自然语言处理系统难以捕捉足够精细的生物医学概念。我们开发了一种针对免疫学与感染病语境的专用命名实体识别(NER)模型。通过与两位临床专家合作,构建并人工标注了371份病例报告数据集,定义了涵盖免疫介导疾病、感染性疾病及相关症状与临床描述的十二类实体。评估了多种建模策略,包括MedicalNER架构结合多类医疗嵌入、基于BERT的词元分类模型以及零样本NER系统。表现最佳的是在临床领域嵌入上训练的Transformer模型,其F1得分为0.89,持续优于基线与零样本方法。专业嵌入与专家标注的结合对捕捉细微疾病术语及提升跨异构生物医学文本的泛化能力尤为关键。提示型大模型基线在相同评估协议下表现显著较差,反映出即使经过详细提示,仍难以生成一致的实体边界输出。该模型为病例报告的结构化分析提供支持,可助力队列识别、疾病监测与临床决策等下游任务。
原文摘要 · Abstract (English)
Extracting detailed clinical information from free-text medical narratives remains a practical challenge for researchers and healthcare systems. Terminology for immune-mediated and infectious diseases is especially inconsistent across sources, which often limits the ability of general-purpose Natural Language Processing (NLP) systems to capture the relevant biomedical concepts with sufficient granularity. We developed a domain-specific Named Entity Recognition (NER) model tailored to identify disease-related entities occurring in immunology and infectious disease contexts. We assembled and manually annotated a dataset of 371 case reports in collaboration with two clinical specialists, defining twelve entity classes covering immune-mediated and infectious conditions as well as related symptoms and clinical descriptors. We evaluated several modeling strategies, including the MedicalNER architecture with multiple healthcare-specific embeddings, a BERT-based token classification model, and zero-shot NER systems. The strongest performance was obtained with a transformer-based model trained on clinical-domain embeddings, which reached an F1 score of 0.89, consistently outperforming baseline and zero-shot approaches. The combination of specialized embeddings and expert annotation proved particularly valuable for capturing nuanced disease terminology and improving generalization across heterogeneous biomedical text. The prompted LLM baseline achieved substantially lower performance under the same evaluation protocol, reflecting difficulties in producing span-consistent outputs for fine-grained entity boundaries despite detailed prompting. The resulting model provides a structured way to analyze case reports and can support downstream tasks such as cohort identification, disease monitoring, and clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。