RelCAT提升病历文本中临床关系抽取的准确率
RelCAT: Advancing Extraction of Clinical Inter-Entity Relationships from Unstructured Electronic Health Records
- 基于MedCAT框架构建交互式工具,融合BERT与Llama模型
- 在n2c2数据集上达0.977宏平均F1,优于现有最佳水平
- 适合医疗信息抽取研究者及临床数据分析团队使用
本研究提出RelCAT(关系概念标注工具包),一个交互式工具、库和工作流,用于分类从临床叙述中提取的实体间关系。基于CogStack MedCAT框架,RelCAT解决了分散在文本中的完整临床关系难以捕捉的问题。该工具包实现先进的机器学习模型如BERT和Llama,并结合成熟的数据评估与训练方法。我们展示了基于MedCATTrainer构建的数据标注工具、模型训练流程,并在公开可用的标准数据集和真实世界英国国家医疗服务体系(NHS)医院临床数据集上评估了该方法。通过大量实验与多种公开模型的对比分析,我们对不同微调策略进行了验证。最终在标准n2c2数据集上取得0.977的宏平均F1分数,超越先前最先进性能;在自收集的NHS数据集上也达到≥0.93的F1分数。
原文摘要 · Abstract (English)
This study introduces RelCAT (Relation Concept Annotation Toolkit), an interactive tool, library, and workflow designed to classify relations between entities extracted from clinical narratives. Building upon the CogStack MedCAT framework, RelCAT addresses the challenge of capturing complete clinical relations dispersed within text. The toolkit implements state-of-the-art machine learning models such as BERT and Llama along with proven evaluation and training methods. We demonstrate a dataset annotation tool (built within MedCATTrainer), model training, and evaluate our methodology on both openly available gold-standard and real-world UK National Health Service (NHS) hospital clinical datasets. We perform extensive experimentation and a comparative analysis of the various publicly available models with varied approaches selected for model fine-tuning. Finally, we achieve macro F1-scores of 0.977 on the gold-standard n2c2, surpassing the previous state-of-the-art performance, and achieve performance of >=0.93 F1 on our NHS gathered datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。