通过细粒度因果干预,消除视觉问答中的语言偏差
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
- 将语言偏差分解为上下文与关键词两类,分别处理
- 在多个VQA模型上提升性能,有效降低偏差影响
- 适合关注模型公平性与多模态对齐的研究者
尽管视觉问答(VQA)取得了显著进展,但由文本信息引入的语言偏差问题仍未解决。以往方法仅从粗粒度角度捕捉偏差,忽略了句子中上下文和关键词等细粒度信息带来的差异性偏差。由于忽视了这些细节,现有方法难以充分建模语言偏差。本文提出一种新的因果干预训练方案CIBi,从细粒度视角消除语言偏差。具体地,将语言偏差分解为上下文偏差与关键词偏差,利用因果干预与对比学习消除上下文偏差并增强多模态表征;同时设计基于反事实生成的问题仅分支,蒸馏并消除关键词偏差。实验表明,CIBi可适配多种VQA模型,取得具有竞争力的性能。
原文摘要 · Abstract (English)
Despite the remarkable advancements in Visual Question Answering (VQA), the challenge of mitigating the language bias introduced by textual information remains unresolved. Previous approaches capture language bias from a coarse-grained perspective. However, the finer-grained information within a sentence, such as context and keywords, can result in different biases. Due to the ignorance of fine-grained information, most existing methods fail to sufficiently capture language bias. In this paper, we propose a novel causal intervention training scheme named CIBi to eliminate language bias from a finer-grained perspective. Specifically, we divide the language bias into context bias and keyword bias. We employ causal intervention and contrastive learning to eliminate context bias and improve the multi-modal representation. Additionally, we design a new question-only branch based on counterfactual generation to distill and eliminate keyword bias. Experimental results illustrate that CIBi is applicable to various VQA models, yielding competitive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。