发现脑胶质瘤生存预测中多模态融合效果来自信号互补而非协同交互。
Quantifying Cross-Modal Interactions in Multimodal Glioma Survival Prediction via InterSHAP: Evidence for Additive Signal Integration
- 用改进的InterSHAP量化多模态交互强度,适配生存分析模型
- 性能越好的模型,跨模态交互越弱(3.0% vs 4.8%)
- 适合关注模型可解释性与联邦学习部署的研究者
针对癌症预后预测中多模态深度学习常假设存在协同效应,但该假设尚未在生存预测中直接验证。本文将基于Shapley交互指数的InterSHAP方法从分类任务扩展至Cox比例风险模型,并应用于胶质瘤生存预测。基于TCGA-GBM与TCGA-LGG数据集(n=575),评估四种结合全切片图像(WSI)与RNA-seq特征的融合架构。核心发现为:预测性能越优(C-index从0.64提升至0.82),测得的跨模态交互强度反而越低(4.8%降至3.0%)。方差分解显示所有架构中贡献稳定:WSI约40%,RNA约55%,交互仅约4%,表明性能提升源于互补信号的叠加而非学习到的协同作用。该研究提供了一种实用的模型审计工具,重新审视了融合架构复杂性的意义,对隐私保护的联邦部署具有启示。
原文摘要 · Abstract (English)
Multimodal deep learning for cancer prognosis is commonly assumed to benefit from synergistic cross-modal interactions, yet this assumption has not been directly tested in survival prediction settings. This work adapts InterSHAP, a Shapley interaction index-based metric, from classification to Cox proportional hazards models and applies it to quantify cross-modal interactions in glioma survival prediction. Using TCGA-GBM and TCGA-LGG data (n=575), we evaluate four fusion architectures combining whole-slide image (WSI) and RNA-seq features. Our central finding is an inverse relationship between predictive performance and measured interaction: architectures achieving superior discrimination (C-index 0.64$\to$0.82) exhibit equivalent or lower cross-modal interaction (4.8\%$\to$3.0\%). Variance decomposition reveals stable additive contributions across all architectures (WSI${\approx}$40\%, RNA${\approx}$55\%, Interaction${\approx}$4\%), indicating that performance gains arise from complementary signal aggregation rather than learned synergy. These findings provide a practical model auditing tool for comparing fusion strategies, reframe the role of architectural complexity in multimodal fusion, and have implications for privacy-preserving federated deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。