arXiv:2506.17903cs.CVcs.AI2025-06IJCAI被引 5

针对医学视觉问答中的语言偏见,提出因果驱动优化框架提升模型鲁棒性。

Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases

  • 从因果与效应双角度设计优化机制,缓解答案偏好与数据不平衡问题。
  • 在多个基准上超越现有方法,尤其在敏感评测集上性能显著提升。
  • 适合医疗AI研究者、希望提升模型公平性的开发者参考。

现有医学视觉问答(Med-VQA)模型常受语言偏见影响,即问题类型与答案类别之间存在虚假关联。为此,本文提出一种新的因果-效应驱动优化框架CEDO,融合三种成熟机制:模态驱动的异构优化(MHO)、梯度引导的模态协同(GMS)和分布自适应损失重缩放(DLR),从因果与效应两方面系统缓解语言偏见。具体而言,MHO采用针对不同模态的自适应学习率,实现异构优化,增强推理鲁棒性;GMS利用帕累托优化促进模态间协同,并强制梯度正交以消除偏见更新,从效应侧缓解捷径偏见;DLR则为各损失分配自适应权重,确保所有答案类别的均衡学习,有效缓解数据集内部的不平衡偏见。在多个传统及偏见敏感基准上的大量实验一致表明,CEDO在鲁棒性上优于当前最优方法。

原文摘要 · Abstract (English)

Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors.

医学视觉问答语言偏见因果优化多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。