提出因果去偏方法,提升视觉常识推理模型的公平性与泛化能力
Causal Debiasing for Visual Commonsense Reasoning
- 基于因果图分析识别文本与图像中的共现及统计偏差
- 构建VCR-OOD数据集,验证模型在跨模态下的泛化性能
- 通过反向门调整与答案字典消除预测捷径,适合关注模型公平性的研究者
视觉常识推理(VCR)要求基于图像回答问题并提供解释。现有方法虽取得高预测准确率,但常忽视数据集中的偏差且缺乏去偏策略。本文分析发现文本与视觉数据中存在共现和统计偏差。为此,我们构建了VCR-OOD数据集,包含VCR-OOD-QA和VCR-OOD-VA两个子集,用于评估模型在双模态下的泛化能力。进一步分析VCR中的因果图与预测捷径,采用反向门调整方法消除偏差。具体地,基于正确答案集合构建字典以消除预测捷径。实验表明,该去偏方法在多个数据集上均有效。
原文摘要 · Abstract (English)
Visual Commonsense Reasoning (VCR) refers to answering questions and providing explanations based on images. While existing methods achieve high prediction accuracy, they often overlook bias in datasets and lack debiasing strategies. In this paper, our analysis reveals co-occurrence and statistical biases in both textual and visual data. We introduce the VCR-OOD datasets, comprising VCR-OOD-QA and VCR-OOD-VA subsets, which are designed to evaluate the generalization capabilities of models across two modalities. Furthermore, we analyze the causal graphs and prediction shortcuts in VCR and adopt a backdoor adjustment method to remove bias. Specifically, we create a dictionary based on the set of correct answers to eliminate prediction shortcuts. Experiments demonstrate the effectiveness of our debiasing method across different datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。