arXiv:2508.19650cs.CV2025-08被引 3

测试大模型在视频中对位置信息的偏见,发现开源模型常偏爱开头或邻近内容。

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

  • 设计新评测基准,控制上下文长度和提问位置,精准检测位置偏见
  • 评估27个主流模型,多数开源模型存在显著头部或邻近偏好
  • 商业模型如Gemini2.5-Pro表现稳定,适合需要全片理解的任务

大型视频语言模型(LVLM)在视频理解方面取得显著进展,推动了相应评测基准的发展。然而,现有基准多关注整个视频序列的整体性能,忽略了诸如上下文位置偏见这一关键但研究不足的特性。本文提出Video-LevelGauge,一个专门用于系统评估LVLM位置偏见的基准。通过标准化探测器与定制化上下文设置,可灵活控制上下文长度、探测位置及上下文类型,模拟多样化真实场景。引入结合统计分析与形态模式识别的综合分析方法以刻画偏见特征。该基准包含438个手动标注视频,涵盖多种类型,生成1,177道高质量选择题和120道开放题,经验证能有效暴露位置偏见。基于此,我们评估了27个最先进的LVLM,包括商业与开源模型。结果显示,许多领先开源模型存在显著位置偏见,通常表现出对视频开头或邻近内容的偏好;相比之下,商业模型如Gemini2.5-Pro在全视频序列中表现一致且优异。进一步分析上下文长度、上下文变化与模型规模,为缓解偏见和指导模型优化提供了可行建议。

原文摘要 · Abstract (English)

Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks generally assess overall performance across entire video sequences, overlooking nuanced behaviors such as contextual positional bias, a critical yet under-explored aspect of LVLM performance. We present Video-LevelGauge, a dedicated benchmark designed to systematically assess positional bias in LVLMs. We employ standardized probes and customized contextual setups, allowing flexible control over context length, probe position, and contextual types to simulate diverse real-world scenarios. In addition, we introduce a comprehensive analysis method that combines statistical measures with morphological pattern recognition to characterize bias. Our benchmark comprises 438 manually curated videos spanning multiple types, yielding 1,177 high-quality multiple-choice questions and 120 open-ended questions, validated for their effectiveness in exposing positional bias. Based on these, we evaluate 27 state-of-the-art LVLMs, including both commercial and open-source models. Our findings reveal significant positional biases in many leading open-source models, typically exhibiting head or neighbor-content preferences. In contrast, commercial models such as Gemini2.5-Pro show impressive, consistent performance across entire video sequences. Further analyses on context length, context variation, and model scale provide actionable insights for mitigating bias and guiding model enhancement . https://github.com/Cola-any/Video-LevelGauge

视频理解位置偏见大模型评测LVLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。