解决医疗多模态数据缺失带来的因果偏差问题,提升模型泛化能力。
Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities
- 基于反事实干预的因果去偏框架,分离真实因果特征与虚假关联。
- 在真实医院和公开数据集上验证,显著提升缺失模态下的预测性能。
- 适合关注医疗AI可解释性与公平性的研究者及临床应用开发者。
医学多模态表示学习旨在整合异构临床数据以构建统一患者表征,支持预测建模,是医学数据挖掘中的关键挑战。然而,现实医疗数据常因成本、流程或患者个体因素导致模态缺失。现有方法多在原始数据或特征空间中处理缺失,却忽视了数据采集过程本身引入的偏差。本文识别出两类阻碍模型泛化的偏差:缺失性偏差(由模态缺失的非随机模式引起)和分布偏差(由影响观测特征与结果的潜在混杂因子导致)。我们对数据生成过程进行结构化因果分析,提出兼容现有直接预测型多模态学习方法的统一框架。该方法包含两个核心组件:(1) 基于后门调整的缺失性去混淆模块,实现因果干预近似;(2) 双分支神经网络,显式解耦因果特征与伪相关。我们在真实世界公开数据集与院内数据集上进行了评估,验证了方法的有效性与因果洞察力。
原文摘要 · Abstract (English)
Medical multimodal representation learning aims to integrate heterogeneous clinical data into unified patient representations to support predictive modeling, which remains an essential yet challenging task in the medical data mining community. However, real-world medical datasets often suffer from missing modalities due to cost, protocol, or patient-specific constraints. Existing methods primarily address this issue by learning from the available observations in either the raw data space or feature space, but typically neglect the underlying bias introduced by the data acquisition process itself. In this work, we identify two types of biases that hinder model generalization: missingness bias, which results from non-random patterns in modality availability, and distribution bias, which arises from latent confounders that influence both observed features and outcomes. To address these challenges, we perform a structural causal analysis of the data-generating process and propose a unified framework that is compatible with existing direct prediction-based multimodal learning methods. Our method consists of two key components: (1) a missingness deconfounding module that approximates causal intervention based on backdoor adjustment and (2) a dual-branch neural network that explicitly disentangles causal features from spurious correlations. We evaluated our method in real-world public and in-hospital datasets, demonstrating its effectiveness and causal insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。