提出新方法消除视频问答中因选项位置引发的错误偏好。
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models
- 通过分解任务和引入公平性度量,识别模型对选项位置的依赖。
- 校准后模型准确率与F1均值显著提升,盲猜现象明显减少。
- 适合关注视频理解模型真实推理能力的研究者使用。
评估视频语言模型(VLMs)极具挑战性。由于透明性高,多选题问答(MCQA)被广泛用于衡量模型性能。然而,现有基准因选择偏差问题,无法全面反映VLMs的推理能力——模型在训练中习得对特定选项位置的偏好。本文对多种VLM架构在主流视频推理数据集上进行系统分析,识别偏差最严重的位置,并揭示模型响应更多依赖于选项位置等表面线索,而非真实视频理解。通过任务拆解与公平性度量适配,提出后处理校准方法BOLD,以平衡该偏差。实验表明,降低选择偏差不仅改善去偏指标,还同步提升准确率与F1均值。该方法通过抑制“盲猜”,相比现有技术更高效、低成本。本研究首次聚焦于视频-文本大模型中的选择偏差问题。
原文摘要 · Abstract (English)
Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, existing MCQA benchmarks fail to capture the full reasoning capabilities of VLMs due to selection bias, when models disproportionately favor certain answer options based on positional patterns observed during training. In this work, we conduct a comprehensive empirical analysis of several VLM architectures across major datasets designed to assess complex video-focused reasoning. We identify where the bias is most pronounced and demonstrate to what extent model responses reflect genuine understanding of video content and related questions, as opposed to reliance on arbitrary patterns or superficial cues, such as answer position. By decomposing the MCQA task and adapting fairness bias metrics to VLMs, we introduce a post-processing calibration technique BOLD to balance this bias. Our results show that reducing selection bias improves not only debiasing metrics but also overall model performance, including Accuracy and F1 Mean score. Our method, by suppressing "blind guessing", offers a more cost- and time-effective approach to mitigating selection bias compared to existing techniques. This study represents the first focused investigation of selection bias in video-to-text LLM-powered models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。