arXiv:2409.14638cs.CLcs.LG2024-09被引 2

微调大模型自动总结住院病程,提升临床编码效率。

Harmonising the Clinical Melody: Tuning Large Language Models for Hospital Course Summarisation in Clinical Coding

  • 用低秩量化微调Llama3等模型,适配医院病程摘要任务。
  • 在MIMIC-III数据上训练,提升摘要与真实病程匹配度。
  • 为临床编码设计新评估指标,验证实用价值。

电子病历中临床文档数量和复杂性持续增加,给临床编码员带来巨大压力,需从海量文本中提取关键信息进行编码。尽管大语言模型在短文本摘要任务中表现良好,但对完整住院过程的摘要仍缺乏研究。本研究采用量化低秩适应(QLoRA)微调三种预训练模型:Llama 3、BioMistral 和 Mistral Instruct v0.1,用于住院病程摘要任务。基于MIMIC-III数据集,通过拼接各类临床记录生成输入文本,并以出院小结中的标准“简要住院过程”作为标注真值。使用BERTScore和ROUGE评估模型性能,同时引入专为临床编码设计的新评估指标验证实际应用价值。结果表明,领域微调可显著提升模型在住院病程摘要上的表现,具备作为临床编码辅助工具的潜力。未来工作应优化数据构建方法,开发更高质量的病程摘要数据集,并尝试更多先进开源模型以逼近闭源模型效果。

原文摘要 · Abstract (English)

The increasing volume and complexity of clinical documentation in Electronic Medical Records systems pose significant challenges for clinical coders, who must mentally process and summarise vast amounts of clinical text to extract essential information needed for coding tasks. While large language models have been successfully applied to shorter summarisation tasks in recent years, the challenge of summarising a hospital course remains an open area for further research and development. In this study, we adapted three pre trained LLMs, Llama 3, BioMistral, Mistral Instruct v0.1 for the hospital course summarisation task, using Quantized Low Rank Adaptation fine tuning. We created a free text clinical dataset from MIMIC III data by concatenating various clinical notes as the input clinical text, paired with ground truth Brief Hospital Course sections extracted from the discharge summaries for model training. The fine tuned models were evaluated using BERTScore and ROUGE metrics to assess the effectiveness of clinical domain fine tuning. Additionally, we validated their practical utility using a novel hospital course summary assessment metric specifically tailored for clinical coding. Our findings indicate that fine tuning pre trained LLMs for the clinical domain can significantly enhance their performance in hospital course summarisation and suggest their potential as assistive tools for clinical coding. Future work should focus on refining data curation methods to create higher quality clinical datasets tailored for hospital course summary tasks and adapting more advanced open source LLMs comparable to proprietary models to further advance this research.

临床编码大模型微调病程摘要医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。