HERBench强制模型融合多段视频证据,检验真实推理能力
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- 设计需整合至少三段不重叠视频线索的题目,杜绝单线索捷径
- 13个顶尖视频大模型平均准确率仅31%-42%,远低于随机猜测
- 揭示检索与融合双重瓶颈,适合研究鲁棒视频理解的学者
视频大语言模型(Video-LLMs)发展迅速,但现有视频问答(VideoQA)基准常允许单线索捷径,难以测试跨时间证据整合能力。我们提出HERBench,一个迫使多证据整合的基准:每个问题至少需要三个来自不同视频片段的非重叠线索。HERBench包含26,806道五选一选择题,涵盖12种组合任务。为量化证据需求,我们引入最小必要帧集(MRFS),即正确回答所需融合的最少帧数,并证明其证据需求高于以往基准。评估13个先进Video-LLM后发现,准确率仅为31%-42%,仅略高于20%的随机基线。我们拆解失败原因:(1) 检索缺陷——帧选择器忽略关键证据;(2) 融合缺陷——即使提供全部证据,模型仍无法有效整合。HERBench因此成为研究鲁棒多证据视频理解的可靠基准。
原文摘要 · Abstract (English)
Video Large Language Models (Video-LLMs) are improving rapidly, yet current Video Question Answering (VideoQA) benchmarks often admit single-cue shortcuts, under-testing reasoning that must integrate evidence across time. We introduce HERBench, a benchmark designed to make multi-evidence integration unavoidable: each question requires at least three non-overlapping cues drawn from distinct video segments. HERBench contains 26,806 five-way multiple-choice questions across 12 compositional tasks. To make evidential demand measurable, we introduce the Minimum Required Frame-Set (MRFS), the smallest number of frames a model must fuse to answer correctly, and show that HERBench imposes higher evidential demand than prior benchmarks. Evaluating 13 state-of-the-art Video-LLMs yields only 31-42% accuracy, only modestly above the 20\% random-guess baseline. We disentangle this failure into two critical bottlenecks: (1) a retrieval deficit, where frame selectors overlook key evidence, and (2) a fusion deficit, where models fail to integrate information even when all necessary evidence is provided. HERBench thus provides a principled benchmark for studying robust multi-evidence video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。