提出新基准GLIMPSE,测试大模型能否真正理解视频而非只看几帧。
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
- 设计需全程观看视频才能答的11类问题,避免只看关键帧
- 人类准确率达94.82%,最强模型仅66.43%,差距明显
- 适合评估视频理解能力,尤其关注时间推理与整体上下文
现有视频评测常类似图像任务,如‘人物在视频中做什么动作’或‘女人连衣裙颜色’,模型只需扫描少数关键帧即可作答,难以检验大视觉语言模型(LVLMs)是否真正具备视频深度推理能力。为此,我们提出GLIMPSE基准,专门评估模型是否能真正‘思考视频’。该基准包含3,269个视频和超过4,342个高度视觉导向的问题,覆盖轨迹分析、时间推理、取证检测等11个类别。所有问题均由人工精心设计,要求完整观看视频并进行全时序推理,无法通过选择性采帧或仅依赖文本回答。人类在该基准上达到94.82%准确率,而表现最佳的GPT-o3模型仅达66.43%,表明当前LVLM仍严重依赖表面特征,缺乏真正的视频理解能力。
原文摘要 · Abstract (English)
Existing video benchmarks often resemble image-based benchmarks, with question types like "What actions does the person perform throughout the video?" or "What color is the woman's dress in the video?" For these, models can often answer by scanning just a few key frames, without deep temporal reasoning. This limits our ability to assess whether large vision-language models (LVLMs) can truly think with videos rather than perform superficial frame-level analysis. To address this, we introduce GLIMPSE, a benchmark specifically designed to evaluate whether LVLMs can genuinely think with videos. Unlike prior benchmarks, GLIMPSE emphasizes comprehensive video understanding beyond static image cues. It consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. All questions are carefully crafted by human annotators and require watching the entire video and reasoning over full video context-this is what we mean by thinking with video. These questions cannot be answered by scanning selected frames or relying on text alone. In human evaluations, GLIMPSE achieves 94.82% accuracy, but current LVLMs face significant challenges. Even the best-performing model, GPT-o3, reaches only 66.43%, highlighting that LVLMs still struggle to move beyond surface-level reasoning to truly think with videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。