arXiv:2509.01167cs.CVcs.CL2025-09被引 1

发现多数视频问答题不依赖帧选择,提出新基准聚焦真正需时序理解的题目。

TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmark

  • 用帧选择敏感度诊断每题是否依赖关键帧
  • 超八成题目对帧选择不敏感,仅8%~33%真正需时序理解
  • 构建小规模精准基准TempCore,适合测试模型时序推理能力

视觉语言模型只能处理有限视频帧,因此帧选择成为实际需求。但现有视频问答基准是否真需要时序帧选择?还是多数问题即使换帧也能答对?我们提出帧选择敏感度(FSS),通过替换最相关帧为最无关帧,测量模型准确率变化。在六个基准和八种模型上,发现多数样本对帧选择不敏感:仅少数真正敏感。结合语言独立性评分(LIS)分析显示,真正时序敏感的样本仅占8%至33%。我们构建了TempCore,从现有基准中提取出这些时序敏感样本的精简评估集,并将在发表后公开代码与每样本标注。

原文摘要 · Abstract (English)

Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current Video QA benchmarks genuinely require temporal frame selection, or can most questions be answered regardless of which frames are shown? We introduce Frame Selection Sensitivity (FSS), a per-sample diagnostic that measures how much VLM accuracy changes when the most relevant frames are replaced with the least relevant ones. Across six benchmarks and eight VLMs, we find that a large majority of samples are frame-agnostic: only a minority are genuinely sensitive to frame choice. Combining FSS with a Language Independence Score (LIS) reveals that merely 8--33% of samples are Temporally Sensitive. We construct TempCore, compact evaluation subsets that isolate these temporal samples from existing benchmarks, and will release code and per-sample annotations upon publication.

视频问答时序理解模型评测基准构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。