针对多模态临床数据融合难题,提出可适配不同任务的跨模态对齐框架。
Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

- 用领域自适应基础模型分别编码影像与病历数据,通过四种策略对齐特征空间
- 在肺栓塞和心血管疾病预测中,融合方法使风险预测准确率提升1.5%~5.4%
- 对比学习对齐效果最稳定,适合临床部署;注意力机制适用于特定任务优化
从多模态临床数据中准确预测时间至事件(TTE)仍面临模态不平衡和分布偏移挑战。本文提出一种基于基础模型的跨模态表示对齐框架,用于连接CT影像与纵向电子健康记录(EHR)数据,旨在实现任务与机构间的泛化。使用领域专用基础模型独立编码CT与EHR,并通过四种有原则的融合策略——晚期融合、对比对齐、交叉注意力与协同注意力——在共享隐空间中对齐。在大规模多中心队列上评估两个临床差异显著的TTE任务:肺栓塞(PE)死亡率与心血管疾病(CVD)结局(PE:训练集3,099例;内部验证1,098例;外部验证435例;CVD:训练集2,951例;内部验证837例;外部验证682例)。当模态贡献相当,融合方法相较单模态基线使一致指数(concordance index)提升1.5%–5.4%。总体而言,对比学习融合(尤其是CLMBR表示)带来最一致且统计显著的提升,尤其在PE死亡率预测中表现突出。对于主要不良心血管事件(MACE),交叉注意力(one-hot)在内部测试中表现最佳,而图像引导的协同注意力在外部验证中表现最优。本研究提出一个可泛化的基础模型驱动跨模态对齐框架,并首次系统分析了模态不平衡下融合行为的影响。结果表明,任务感知的多模态对齐是实现鲁棒泛化与可扩展临床部署的必要设计原则。
原文摘要 · Abstract (English)
Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift. We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize across tasks and institutions. CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion strategies: late fusion, contrastive alignment, cross-attention, and co-attention. We evaluate two clinically distinct TTE tasks: pulmonary embolism (PE) mortality and cardiovascular disease (CVD) outcomes, on large-scale multi-institutional cohorts (PE: N=3,099 train; 1,098 internal; 435 external; CVD: N=2,951 train; 837 internal; 682 external). Fusion consistently improves concordance index by 1.5-5.4% over unimodal baselines when modalities contribute comparably. Overall, contrastive multimodal fusion, particularly with CLMBR representations, provided the most consistent and statistically robust improvements, especially for PE mortality prediction. For MACE, cross-attention (one-hot) achieved the highest internal performance and image-guided co-attention achieved the best external performance. We therefore introduce a generalizable foundation model-based cross-modal alignment framework and provide the first systematic analysis of fusion behavior under modality imbalance in TTE prediction. Our results establish task-aware multimodal alignment as a necessary design principle for robust generalization and scalable clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。