评估开源大模型在医疗文本摘要中的幻觉与关键信息提取能力
Hallucinations and Key Information Extraction in Medical Texts: A Comprehensive Assessment of Open-Source Large Language Models
- 对比多个开源大模型提取出院记录中的关键事件
- 模型对随访建议的识别准确率显著低于入院原因和住院事件
- 揭示医疗摘要中幻觉问题,提醒临床应用需谨慎
临床摘要在医疗中至关重要,能将复杂的医学数据提炼为易于理解的信息,提升患者理解和护理管理效率。大型语言模型(LLMs)凭借其先进的自然语言理解能力,在自动化和提高摘要准确性方面展现出巨大潜力,尤其适用于需要精确简洁信息传递的医疗文本摘要场景。本文系统评估了开源LLMs在从出院报告中提取关键事件(包括入院原因、住院主要事件及重要随访措施)方面的表现,并分析了模型生成摘要中各类幻觉的普遍性。幻觉检测至关重要,因为它直接影响信息可靠性,可能影响患者治疗与预后。通过全面模拟,我们严格评估了模型在临床摘要中的性能,结果表明:尽管如Qwen2.5和DeepSeek-v2等模型在捕捉入院原因和住院事件方面表现良好,但在识别随访建议方面普遍不够一致,暴露出当前大模型在全面临床摘要任务中的深层挑战。
原文摘要 · Abstract (English)
Clinical summarization is crucial in healthcare as it distills complex medical data into digestible information, enhancing patient understanding and care management. Large language models (LLMs) have shown significant potential in automating and improving the accuracy of such summarizations due to their advanced natural language understanding capabilities. These models are particularly applicable in the context of summarizing medical/clinical texts, where precise and concise information transfer is essential. In this paper, we investigate the effectiveness of open-source LLMs in extracting key events from discharge reports, including admission reasons, major in-hospital events, and critical follow-up actions. In addition, we also assess the prevalence of various types of hallucinations in the summaries produced by these models. Detecting hallucinations is vital as it directly influences the reliability of the information, potentially affecting patient care and treatment outcomes. We conduct comprehensive simulations to rigorously evaluate the performance of these models, further probing the accuracy and fidelity of the extracted content in clinical summarization. Our results reveal that while the LLMs (e.g., Qwen2.5 and DeepSeek-v2) perform quite well in capturing admission reasons and hospitalization events, they are generally less consistent when it comes to identifying follow-up recommendations, highlighting broader challenges in leveraging LLMs for comprehensive summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。