arXiv:2503.16069cs.CV2025-03中稿 · MICCAI 2025被引 9

提出可解耦的多模态注意力融合模型,提升癌症生存预测准确性与可解释性。

Disentangled and Interpretable Multimodal Attention Fusion for Cancer Survival Prediction

  • 分离模态内与模态间交互,学习独立的共享与特异表征
  • 在4个公开数据集上性能提升1.85%,解耦程度提高23.7%
  • 结合Shapley值分析,揭示多模态数据在癌症生物学中的作用机制

为提升基于全切片图像与转录组数据的癌症生存预测效果,需同时捕捉模态共享与模态特异性信息。然而,现有多模态框架常使表征纠缠,影响可解释性并可能抑制判别特征。为此,我们提出解耦且可解释的多模态注意力融合(DIMAF)框架,通过注意力机制分离模态内与模态间交互,学习独立的模态特异与模态共享表示。引入基于距离相关性的损失以促进两类表征解耦,并结合Shapley加性解释评估其对生存预测的相对贡献。在四个公开癌症生存数据集上,相比当前最优多模态模型,性能平均提升1.85%,解耦程度提高23.7%。除性能提升外,该可解释框架支持对癌症生物学中模态间及模态内交互机制的深入探索。

原文摘要 · Abstract (English)

To improve the prediction of cancer survival using whole-slide images and transcriptomics data, it is crucial to capture both modality-shared and modality-specific information. However, multimodal frameworks often entangle these representations, limiting interpretability and potentially suppressing discriminative features. To address this, we propose Disentangled and Interpretable Multimodal Attention Fusion (DIMAF), a multimodal framework that separates the intra- and inter-modal interactions within an attention-based fusion mechanism to learn distinct modality-specific and modality-shared representations. We introduce a loss based on Distance Correlation to promote disentanglement between these representations and integrate Shapley additive explanations to assess their relative contributions to survival prediction. We evaluate DIMAF on four public cancer survival datasets, achieving a relative average improvement of 1.85% in performance and 23.7% in disentanglement compared to current state-of-the-art multimodal models. Beyond improved performance, our interpretable framework enables a deeper exploration of the underlying interactions between and within modalities in cancer biology.

多模态融合生存预测可解释性癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。