arXiv:2501.05690cs.CVcs.CL2025-01中稿 · ICME2024被引 2

用知识蒸馏削弱视觉问答中的语言偏见,提升模型泛化能力

Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation

  • 通过教师模型的软标签提供正则化与语义引导,抑制对常见答案的依赖
  • 在VQA-CPv2数据集上实现最新性能,显著优于已有方法
  • 适合关注模型公平性与鲁棒性的研究人员使用

以往研究表明,视觉问答(VQA)模型容易依赖语言先验进行预测,导致泛化能力下降。本文提出KDAR方法,利用知识蒸馏缓解此类偏见。通过来自训练良好的教师模型的软标签,对过度拟合常见答案的行为进行惩罚,并提供语义指导以缩小候选答案范围。此外,设计自适应样本权重学习策略,动态调整每个样本的重要性,进一步缓解偏差。实验表明,该方法在分布外(OOD)和分布内(IID)设置下均表现更优,在VQA-CPv2 OOD基准上达到当前最优性能。

原文摘要 · Abstract (English)

Previous studies have pointed out that visual question answering (VQA) models are prone to relying on language priors for answer predictions. In this context, predictions often depend on linguistic shortcuts rather than a comprehensive grasp of multimodal knowledge, which diminishes their generalization ability. In this paper, we propose a novel method, namely, KDAR, leveraging knowledge distillation to address the prior-dependency dilemmas within the VQA task. Specifically, the regularization effect facilitated by soft labels from a well-trained teacher is employed to penalize overfitting to the most common answers. The soft labels, which serve a regularization role, also provide semantic guidance that narrows the range of candidate answers. Additionally, we design an adaptive sample-wise reweighting learning strategy to further mitigate bias by dynamically adjusting the importance of each sample. Experimental results demonstrate that our method enhances performance in both OOD and IID settings. Our method achieves state-of-the-art performance on the VQA-CPv2 out-of-distribution (OOD) benchmark, significantly outperforming previous state-of-the-art approaches.

视觉问答知识蒸馏语言偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。