用多模态融合提升心衰患者短期死亡预测准确率
Is Clinical Text Enough? A Multimodal Study on Mortality Prediction in Heart Failure Patients
- 引入实体级表示增强临床文本,优于传统摘要嵌入
- 文本与结构化数据联合建模效果最佳,AUC达0.83
- 大模型提示工程表现不稳定,文本单独输入更可靠
心力衰竭(HF)短期死亡预测仍具挑战性,尤其依赖结构化电子健康记录(EHR)时。我们在法国HF队列上评估了基于Transformer的模型,对比了仅文本、仅结构化数据、多模态及大语言模型(LLM)方法。结果表明,将实体级表示融入临床文本可显著提升预测性能,优于仅使用CLS嵌入;监督式多模态融合在文本与结构化变量上表现最优,整体AUC达0.83。相反,大语言模型在不同模态与解码策略下表现不一致,仅文本提示反而优于结构化或多模态输入。研究证实,具备实体感知能力的多模态Transformer是短期心衰预后预测最可靠的方案,而当前大模型提示工程在临床决策支持中仍有限。
原文摘要 · Abstract (English)
Accurate short-term mortality prediction in heart failure (HF) remains challenging, particularly when relying on structured electronic health record (EHR) data alone. We evaluate transformer-based models on a French HF cohort, comparing text-only, structured-only, multimodal, and LLM-based approaches. Our results show that enriching clinical text with entity-level representations improves prediction over CLS embeddings alone, and that supervised multimodal fusion of text and structured variables achieves the best overall performance. In contrast, large language models perform inconsistently across modalities and decoding strategies, with text-only prompts outperforming structured or multimodal inputs. These findings highlight that entity-aware multimodal transformers offer the most reliable solution for short-term HF outcome prediction, while current LLM prompting remains limited for clinical decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。