多模态融合框架提升心脏病多任务分析准确率
A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- 动态融合文本、信号、影像数据,自适应权重分配
- 在多个临床任务中超越现有方法,最高提升12.3%
- 适合临床决策支持系统开发与医学人工智能研究
当前心血管管理需整合多模态心脏数据,每种模态提供互补的生理特征。现有方法受限于:1)患者与时间对齐的多模态数据稀缺;2)依赖单一模态或固定组合;3)对齐策略偏重跨模态相似性而非互补性;4)任务范围狭窄。为此,研究构建了包含实验室检验、心电图、超声心动图及临床结局的综合多模态数据集,并提出统一框架TGMM。该框架包含:1)MedFlexFusion模块,捕捉各模态独特性与互补性,动态整合不同来源数据;2)文本引导模块,生成面向诊断、风险分层与信息检索等任务的定制表示;3)响应模块,输出各项任务最终决策。实验表明,TGMM在多项临床任务中优于当前最优方法,且在另一公开数据集上验证了其鲁棒性。
原文摘要 · Abstract (English)
Contemporary cardiovascular management involves complex consideration and integration of multimodal cardiac datasets, where each modality provides distinct but complementary physiological characteristics. While the effective integration of multiple modalities could yield a holistic clinical profile that accurately models the true clinical situation with respect to data modalities and their relatives weightings, current methodologies remain limited by: 1) the scarcity of patient- and time-aligned multimodal data; 2) reliance on isolated single-modality or rigid multimodal input combinations; 3) alignment strategies that prioritize cross-modal similarity over complementarity; and 4) a narrow single-task focus. In response to these limitations, a comprehensive multimodal dataset was curated for immediate application, integrating laboratory test results, electrocardiograms, and echocardiograms with clinical outcomes. Subsequently, a unified framework, Textual Guidance Multimodal fusion for Multiple cardiac tasks (TGMM), was proposed. TGMM incorporated three key components: 1) a MedFlexFusion module designed to capture the unique and complementary characteristics of medical modalities and dynamically integrate data from diverse cardiac sources and their combinations; 2) a textual guidance module to derive task-relevant representations tailored to diverse clinical objectives, including heart disease diagnosis, risk stratification and information retrieval; and 3) a response module to produce final decisions for all these tasks. Furthermore, this study systematically explored key features across multiple modalities and elucidated their synergistic contributions in clinical decision-making. Extensive experiments showed that TGMM outperformed state-of-the-art methods across multiple clinical tasks, with additional validation confirming its robustness on another public dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。