arXiv:2603.26028cs.CV2026-03

让医学视觉问答模型自动剔除数据偏见,专注真实诊断证据。

Learning to Trim: End-to-End Causal Graph Pruning with Dynamic Anatomical Feature Banks for Medical VQA

  • 用动态特征库捕捉常见解剖模式,作为数据偏差的代理
  • 可微剪枝模块软抑制与全局模式强相关的特征,突出个体证据
  • 在多个医学VQA数据集上显著提升模型泛化能力

医学视觉问答(MedVQA)模型常因依赖数据集特有的关联(如重复的解剖结构或问题类型规律)而泛化能力有限,而非真正诊断依据。现有因果方法多为静态调整或事后修正。为此,我们提出可学习的因果剪枝(LCT)框架,将因果剪枝融入端到端优化。引入动态解剖特征库(DAFB),通过动量机制更新,捕捉频繁出现的解剖与语言模式的全局原型,近似数据级规律。设计可微剪枝模块,估计实例表示与全局特征库间的依赖关系:与全局原型高度相关的特征被软抑制,而实例特异性证据被强化。该可学习机制使模型能自适应地优先关注因果信号而非虚假相关。在VQA-RAD、SLAKE、SLAKE-CP和PathVQA上的实验表明,LCT在鲁棒性和泛化性上持续优于现有去偏策略。

原文摘要 · Abstract (English)

Medical Visual Question Answering (MedVQA) models often exhibit limited generalization due to reliance on dataset-specific correlations, such as recurring anatomical patterns or question-type regularities, rather than genuine diagnostic evidence. Existing causal approaches are typically implemented as static adjustments or post-hoc corrections. To address this issue, we propose a Learnable Causal Trimming (LCT) framework that integrates causal pruning into end-to-end optimization. We introduce a Dynamic Anatomical Feature Bank (DAFB), updated via a momentum mechanism, to capture global prototypes of frequent anatomical and linguistic patterns, serving as an approximation of dataset-level regularities. We further design a differentiable trimming module that estimates the dependency between instance-level representations and the global feature bank. Features highly correlated with global prototypes are softly suppressed, while instance-specific evidence is emphasized. This learnable mechanism encourages the model to prioritize causal signals over spurious correlations adaptively. Experiments on VQA-RAD, SLAKE, SLAKE-CP and PathVQA demonstrate that LCT consistently improves robustness and generalization over existing debiasing strategies.

医学VQA因果推理去偏可微剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。