提出因果方法缓解医学图像问答中的模态偏好偏差
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
- 基于因果图构建反事实训练机制,消除模态偏好
- 在新重构数据集上,性能显著优于非因果模型
- 适合关注医疗多模态公平性的研究者和临床辅助系统开发者
医学视觉问答(MedVQA)对提升临床诊断效率至关重要,能及时准确回答医生关于医学影像的问题。现有模型存在模态偏好偏差,即预测过度依赖问题而忽视图像信息,无法有效学习多模态知识。为此,本文提出医学反事实视觉问答(MedCFVQA)模型,通过引入因果图在训练中消除推理时的模态偏好。现有MedVQA数据集存在问题与答案间的强先验依赖,导致模型即使存在偏差也能取得较好表现。为此,我们利用现有数据集重构新数据集,改变训练与测试集中问题与答案之间的先验依赖关系(CP)。大量实验表明,MedCFVQA在SLAKE、RadVQA及对应的SLAKE-CP、RadVQA-CP数据集上均显著优于非因果基线模型。
原文摘要 · Abstract (English)
Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses to clinicians' inquiries regarding medical images. Existing MedVQA models suffered from modality preference bias, where predictions are heavily dominated by one modality while overlooking the other (in MedVQA, usually questions dominate the answer but images are overlooked), thereby failing to learn multimodal knowledge. To overcome the modality preference bias, we proposed a Medical CounterFactual VQA (MedCFVQA) model, which trains with bias and leverages causal graphs to eliminate the modality preference bias during inference. Existing MedVQA datasets exhibit substantial prior dependencies between questions and answers, which results in acceptable performance even if the model significantly suffers from the modality preference bias. To address this issue, we reconstructed new datasets by leveraging existing MedVQA datasets and Changed their P3rior dependencies (CP) between questions and their answers in the training and test set. Extensive experiments demonstrate that MedCFVQA significantly outperforms its non-causal counterpart on both SLAKE, RadVQA and SLAKE-CP, RadVQA-CP datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。