arXiv:2604.20306cs.CVcs.AI2026-04被引 1

提出双因果推理框架,提升医学视觉问答的可信诊断能力。

Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA

论文配图:Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA
图 1 · 摘自论文原文
  • 融合后门调整与工具变量学习,同时处理可观测和不可观测混杂因素。
  • 在四个数据集上显著优于现有方法,尤其在分布外泛化上提升明显。
  • 适合关注医疗AI可解释性与鲁棒性的研究者和临床应用开发者。

医学视觉问答(MedVQA)旨在基于复杂医学图像和问题生成具有临床意义的答案。然而,现有方法常过度拟合于跨模态表面相关性,忽视多模态医学数据中的内在偏差,导致模型易受跨模态混杂影响,严重削弱其可信诊断推理能力。为此,本文提出一种新颖的双因果推理(DCI)框架,首次将后门调整(BDA)与工具变量(IV)学习统一建模,以同时应对可观测和不可观测混杂因素。具体而言,构建结构因果模型(SCM),通过BDA缓解可观测的跨模态偏差(如常见视觉-文本共现),并通过从共享潜在空间学习的IV补偿不可观测混杂因子。为确保IV有效性,设计互信息约束,最大化其与融合多模态表示的相关性,同时最小化其与不可观测混杂因子及目标答案的关联。该双机制使模型提取去混杂表征,捕捉真实因果关系。在四个基准数据集(SLAKE、SLAKE-CP、VQA-RAD、PathVQA)上的实验表明,本方法持续优于现有方法,尤其在分布外(OOD)泛化上表现突出。定性分析进一步证实,DCI显著增强跨模态推理的可解释性与鲁棒性,明确分离真实因果效应与虚假跨模态捷径。

原文摘要 · Abstract (English)

Medical Visual Question Answering (MedVQA) aims to generate clinically reliable answers conditioned on complex medical images and questions. However, existing methods often overfit to superficial cross-modal correlations, neglecting the intrinsic biases embedded in multimodal medical data. Consequently, models become vulnerable to cross-modal confounding effects, severely hindering their ability to provide trustworthy diagnostic reasoning. To address this limitation, we propose a novel Dual Causal Inference (DCI) framework for MedVQA. To the best of our knowledge, DCI is the first unified architecture that integrates Backdoor Adjustment (BDA) and Instrumental Variable (IV) learning to jointly tackle both observable and unobserved confounders. Specifically, we formulate a Structural Causal Model (SCM) where observable cross-modal biases (e.g., frequent visual and textual co-occurrences) are mitigated via BDA, while unobserved confounders are compensated using an IV learned from a shared latent space. To guarantee the validity of the IV, we design mutual information constraints that maximize its dependence on the fused multimodal representations while minimizing its associations with the unobserved confounders and target answers. Through this dual mechanism, DCI extracts deconfounded representations that capture genuine causal relationships. Extensive experiments on four benchmark datasets, SLAKE, SLAKE-CP, VQA-RAD, and PathVQA, demonstrate that our method consistently outperforms existing approaches, particularly in out-of-distribution (OOD) generalization. Furthermore, qualitative analyses confirm that DCI significantly enhances the interpretability and robustness of cross-modal reasoning by explicitly disentangling true causal effects from spurious cross-modal shortcuts.

医学问答因果推理多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。