解决医疗多模态数据中的系统性偏见,提升模型泛化能力。
Robust Multimodal Representation Learning in Healthcare
- 通过因果分析识别潜在混杂因子带来的偏见
- 双流网络分离因果特征与虚假相关,提升预测性能
- 可嵌入现有方法,适合医疗多模态研究者使用
医疗多模态表征学习旨在整合异构数据生成统一的患者表征以支持临床结果预测。然而,真实医疗数据常因多重来源产生系统性偏见,严重影响模型泛化能力。现有方法多关注多模态融合,忽视了影响泛化能力的内在偏见特征。为此,本文提出一种双流特征去相关框架,通过潜在混杂因子引入的结构因果分析识别并处理偏见。该方法采用因果偏见去相关机制,结合双流神经网络,利用广义交叉熵损失和互信息最小化实现有效去相关。框架具有模型无关性,可集成至现有医疗多模态学习方法中。在MIMIC-IV、eICU和ADNI数据集上的全面实验表明,该方法实现一致的性能提升。
原文摘要 · Abstract (English)
Medical multimodal representation learning aims to integrate heterogeneous data into unified patient representations to support clinical outcome prediction. However, real-world medical datasets commonly contain systematic biases from multiple sources, which poses significant challenges for medical multimodal representation learning. Existing approaches typically focus on effective multimodal fusion, neglecting inherent biased features that affect the generalization ability. To address these challenges, we propose a Dual-Stream Feature Decorrelation Framework that identifies and handles the biases through structural causal analysis introduced by latent confounders. Our method employs a causal-biased decorrelation framework with dual-stream neural networks to disentangle causal features from spurious correlations, utilizing generalized cross-entropy loss and mutual information minimization for effective decorrelation. The framework is model-agnostic and can be integrated into existing medical multimodal learning methods. Comprehensive experiments on MIMIC-IV, eICU, and ADNI datasets demonstrate consistent performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。