小模型视频理解性能受采帧策略影响显著,首个精准控制采帧的基准测试揭示此偏差。
Frame Sampling Strategies Matter: A Benchmark for small vision language models
- 构建首个精确控制采帧策略的小型视觉语言模型视频评测基准
- 发现不同采帧方式导致模型表现差异超30%,存在显著采样偏差
- 适合关注视频理解公平评测与模型可复现性的研究者
在视频任务中评估视觉语言模型(VLMs)极为复杂,因性能同时受模型视觉表征能力与输入构建所用帧采样策略的影响。当前视频基准测试可能因采用不同采帧方式而产生严重偏差。本文提出首个针对先进小型视觉语言模型(SVLMs)的帧级精确视频问答基准,所有模型在受控的采帧策略下进行评估。结果证实了这一偏差的存在,并揭示了小模型在不同采帧技术下的数据特异性和任务特异性行为。通过开源评测代码,我们为社区提供可复现、无偏的视频VLM评估协议,并强调未来研究需为每个数据集设计标准化的采帧策略。
原文摘要 · Abstract (English)
Comparing vision language models on videos is particularly complex, as the performances is jointly determined by the model's visual representation capacity and the frame-sampling strategy used to construct the input. Current video benchmarks are suspected to suffer from substantial frame-sampling bias, as models are evaluated with different frame selection strategies. In this work, we propose the first frame-accurate benchmark of state-of-the-art small VLMs for video question-answering, evaluated under controlled frame-sampling strategies. Our results confirm the suspected bias and highlight both data-specific and task-specific behaviors of SVLMs under different frame-sampling techniques. By open-sourcing our benchmarking code, we provide the community with a reproducible and unbiased protocol for evaluating video VLMs and emphasize the need for standardized frame-sampling strategies tailored to each benchmarking dataset in future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。