评测大模型对长期病历的时间推理能力,发现其仍难把握疾病演变和罕见病预测。
Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction
- 用多模态电子病历数据重构任务,测试大模型在长时序下的摘要与预测能力。
- 长上下文提升信息整合,但未显著改善时间推理与罕见病预测性能。
- 适合关注医疗AI、时间序列建模的科研人员参考。
大语言模型在临床文本摘要方面展现出潜力,但对其处理跨时间分布的多模态患者轨迹的能力仍研究不足。本研究系统评估了多种先进开源LLM及其检索增强生成(RAG)变体和思维链(CoT)提示方法,在长上下文临床摘要与预测任务中的表现。通过重构现有任务,包括从两个公开EHR数据集进行出院摘要和诊断预测,考察模型融合结构化与非结构化病历数据并进行时间连贯性推理的能力。结果表明,长上下文窗口虽提升了输入整合效果,但并未一致增强临床推理能力;模型在时间进程理解和罕见病预测方面仍存困难。尽管RAG在某些情况下减少了幻觉,但未能根本解决上述局限。本工作填补了长期临床文本摘要的研究空白,为基于多模态数据和时间推理的大模型评估建立了基础。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have shown potential in clinical text summarization, but their ability to handle long patient trajectories with multi-modal data spread across time remains underexplored. This study systematically evaluates several state-of-the-art open-source LLMs, their Retrieval Augmented Generation (RAG) variants and chain-of-thought (CoT) prompting on long-context clinical summarization and prediction. We examine their ability to synthesize structured and unstructured Electronic Health Records (EHR) data while reasoning over temporal coherence, by re-engineering existing tasks, including discharge summarization and diagnosis prediction from two publicly available EHR datasets. Our results indicate that long context windows improve input integration but do not consistently enhance clinical reasoning, and LLMs are still struggling with temporal progression and rare disease prediction. While RAG shows improvements in hallucination in some cases, it does not fully address these limitations. Our work fills the gap in long clinical text summarization, establishing a foundation for evaluating LLMs with multi-modal data and temporal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。