发现视频模型过度依赖语言,提出新评测与解法
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
- 构建视频语言偏见评测基准,识别模型偏倚问题
- 现有模型在视频任务中普遍受语言干扰,准确率下降
- 提出多分支对比解码,无需重训即可缓解偏见
近期大型视觉语言模型(LVLMs)在多模态任务中取得显著进展,但本文揭示了一个被忽视的问题:视频相关LVLM存在语言偏见,即模型更依赖语言信息而非视频内容,导致错误回答。为此,我们构建了视频语言偏见评测基准(VidLBEval),包含模糊视频对比和疑问句探测两类任务,设计相应评估指标以惩罚语言偏倚。同时提出多分支对比解码(MCD)方法,引入两个专家分支,协同对抗仅文本分支可能引发的偏见。实验表明:现有视频型LVLM(包括闭源与开源模型)普遍存在语言偏见问题;所提MCD能有效缓解该问题,且无需额外训练或修改模型结构即可保持通用能力。
原文摘要 · Abstract (English)
Recently, Large Vision-Language Models (LVLMs) have made significant strides across diverse multimodal tasks and benchmarks. This paper reveals a largely under-explored problem from existing video-involved LVLMs - language bias, where models tend to prioritize language over video and thus result in incorrect responses. To address this research gap, we first collect a Video Language Bias Evaluation Benchmark, which is specifically designed to assess the language bias in video-involved LVLMs through two key tasks: ambiguous video contrast and interrogative question probing. Accordingly, we design accompanied evaluation metrics that aim to penalize LVLMs being biased by language. In addition, we also propose Multi-branch Contrastive Decoding (MCD), introducing two expert branches to simultaneously counteract language bias potentially generated by the amateur text-only branch. Our experiments demonstrate that i) existing video-involved LVLMs, including both proprietary and open-sourced, are largely limited by the language bias problem; ii) our MCD can effectively mitigate this issue and maintain general-purpose capabilities in various video-involved LVLMs without any additional retraining or alteration to model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。