arXiv:2411.03805cs.CL2024-11被引 29

对比大模型生成肺癌出院小结,发现LLaMA 3表现稳健且简洁。

A Comparative Study of Recent Large Language Models on Generating Hospital Discharge Summaries for Lung Cancer Patients

  • 用1099例肺癌患者病历测试GPT-4o、LLaMA 3等模型生成摘要
  • GPT-4o与微调后的LLaMA 3在指标上领先,尤其语义相似度高
  • LLaMA 3在不同长度病历下保持摘要简洁,适合临床稳定使用

生成出院小结是临床实践中的关键但耗时的任务,对传递患者信息和保障医疗连续性至关重要。近年来大语言模型(LLMs)在理解与总结复杂医学文本方面取得显著进展。本研究旨在探索LLMs如何减轻人工总结负担、提升工作流效率并支持临床决策。基于1,099例肺癌患者病历,其中50例用于测试,102例用于模型微调。评估了GPT-3.5、GPT-4、GPT-4o和LLaMA 3 8b在生成出院小结方面的表现,采用词元级指标(BLEU、ROUGE-1、ROUGE-2、ROUGE-L)及生成摘要与医生撰写标准之间的语义相似度评分。结果显示各模型总结能力差异显著:GPT-4o与微调后的LLaMA 3在词元级指标上表现最优;而LLaMA 3在不同长度输入下均保持摘要简洁。语义相似度分析表明,GPT-4o与LLaMA 3最能捕捉临床相关性。研究揭示了LLMs在生成出院小结中的有效性,突显了LLaMA 3在多种临床情境下维持清晰与相关性的鲁棒性。结果表明自动化摘要工具有望提升文档精确度与效率,最终改善患者照护与医疗机构运营能力。

原文摘要 · Abstract (English)

Generating discharge summaries is a crucial yet time-consuming task in clinical practice, essential for conveying pertinent patient information and facilitating continuity of care. Recent advancements in large language models (LLMs) have significantly enhanced their capability in understanding and summarizing complex medical texts. This research aims to explore how LLMs can alleviate the burden of manual summarization, streamline workflow efficiencies, and support informed decision-making in healthcare settings. Clinical notes from a cohort of 1,099 lung cancer patients were utilized, with a subset of 50 patients for testing purposes, and 102 patients used for model fine-tuning. This study evaluates the performance of multiple LLMs, including GPT-3.5, GPT-4, GPT-4o, and LLaMA 3 8b, in generating discharge summaries. Evaluation metrics included token-level analysis (BLEU, ROUGE-1, ROUGE-2, ROUGE-L) and semantic similarity scores between model-generated summaries and physician-written gold standards. LLaMA 3 8b was further tested on clinical notes of varying lengths to examine the stability of its performance. The study found notable variations in summarization capabilities among LLMs. GPT-4o and fine-tuned LLaMA 3 demonstrated superior token-level evaluation metrics, while LLaMA 3 consistently produced concise summaries across different input lengths. Semantic similarity scores indicated GPT-4o and LLaMA 3 as leading models in capturing clinical relevance. This study contributes insights into the efficacy of LLMs for generating discharge summaries, highlighting LLaMA 3's robust performance in maintaining clarity and relevance across varying clinical contexts. These findings underscore the potential of automated summarization tools to enhance documentation precision and efficiency, ultimately improving patient care and operational capability in healthcare settings.

医疗AI大模型出院小结LLaMA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。