用大模型自动提取病历中的关键信息,提升电子病历结构化效率。
Extracting Patient History from Clinical Text: A Comparative Study of Clinical Large Language Models
- 微调临床大模型识别患者主诉、现病史等医学实体
- 微调后的模型比GPT-4零样本提升性能,尤其在长实体上表现更好
- 适合需要高效处理临床文本的医疗系统开发者
从自由文本临床记录中提取与主诉(CC)、现病史(HPI)及既往史、家族史、社会史(PFSH)相关的医学历史实体(MHEs),有助于将非结构化病历转化为标准化电子健康记录(EHR),支持连续性照护、医疗编码和质量评估等下游任务。通过本地部署微调的临床大语言模型(cLLMs)可保护敏感数据。本研究评估了七种先进cLLMs在识别相关MHEs上的表现,并分析了病历特征对准确率的影响。我们在MTSamples数据集的61份门诊病历中标注了1,449个MHEs,对cLLMs进行微调,并测试了引入问题、检查、治疗等基础医学实体(BMEs)后模型的表现。同时以GPT-4o零样本作为对比。结果显示,微调后的模型可减少20%以上的提取耗时。尽管如此,由于实体语义多样性和非医学词汇频繁出现,多数类型仍难识别。其中训练最充分的GatorTron和GatorTronS表现最佳;整合预识别的BME信息提升了部分实体的识别效果。错误分析表明:长实体更难识别,病历长度与错误率无显著关联,而带标题的清晰段落结构有助于提高准确性。
原文摘要 · Abstract (English)
Extracting medical history entities (MHEs) related to a patient's chief complaint (CC), history of present illness (HPI), and past, family, and social history (PFSH) helps structure free-text clinical notes into standardized EHRs, streamlining downstream tasks like continuity of care, medical coding, and quality metrics. Fine-tuned clinical large language models (cLLMs) can assist in this process while ensuring the protection of sensitive data via on-premises deployment. This study evaluates the performance of cLLMs in recognizing CC/HPI/PFSH-related MHEs and examines how note characteristics impact model accuracy. We annotated 1,449 MHEs across 61 outpatient-related clinical notes from the MTSamples repository. To recognize these entities, we fine-tuned seven state-of-the-art cLLMs. Additionally, we assessed the models' performance when enhanced by integrating, problems, tests, treatments, and other basic medical entities (BMEs). We compared the performance of these models against GPT-4o in a zero-shot setting. To further understand the textual characteristics affecting model accuracy, we conducted an error analysis focused on note length, entity length, and segmentation. The cLLMs showed potential in reducing the time required for extracting MHEs by over 20%. However, detecting many types of MHEs remained challenging due to their polysemous nature and the frequent involvement of non-medical vocabulary. Fine-tuned GatorTron and GatorTronS, two of the most extensively trained cLLMs, demonstrated the highest performance. Integrating pre-identified BME information improved model performance for certain entities. Regarding the impact of textual characteristics on model performance, we found that longer entities were harder to identify, note length did not correlate with a higher error rate, and well-organized segments with headings are beneficial for the extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。