arXiv:2411.09105cs.CVcs.AI2024-11被引 4

构建可控视频认知评估基准,揭示大模型在抽象思维上的短板。

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

  • 用程序化生成合成视频,精确控制视觉与任务难度
  • 顶尖模型在抽象任务上平均仅48.8%正确率,复杂度上升时性能降15%
  • 适合研究视频理解、认知建模与模型可解释性的学者

近年来,大型视频语言模型(LVLMs)在多模态视频理解方面取得显著进展。然而,这些模型是否具备高阶任务所需的认知能力,尤其是符号与抽象感知能力,仍不明确。现有基准多依赖真实世界标注视频,难以控制内容与难度,诊断能力有限。为此,我们提出VideoCogQA,一个受游戏环境启发的可扩展、全可控基准,用于评估LVLM的认知能力。通过程序化引擎生成合成视频,可精细调控视觉元素、时间动态与任务难度,实现对视频认知能力的聚焦评估,不受视觉场景语义先验影响。数据集包含800个视频和3,280个问答对,涵盖抽象概念、符号元素与多模态融合任务,具有不同难度层级。实验表明,即使最先进的模型如GPT-4o,在抽象概念任务上平均准确率仅为48.8%,且随着任务复杂度增加,性能下降15%,凸显了当前模型在保持稳定表现方面的挑战。本工作旨在揭示现有LVLM的局限性,并为未来更贴近人类认知过程的模型设计提供启示。

原文摘要 · Abstract (English)

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks, particularly those involving symbolic and abstract perception. Existing benchmarks typically rely on real-world, annotated videos, which lack control over video content and inherent difficulty, limiting their diagnostic power. To bridge this gap, we propose VideoCogQA, a scalable and fully controllable benchmark inspired by game-world environments, designed to evaluate the cognitive abilities of LVLMs. By generating synthetic videos via a programmatic engine, VideoCogQA allows fine-grained control over visual elements, temporal dynamics, and task difficulty. This approach enables a focused evaluation of video cognitive abilities, independent of prior knowledge from visual scene semantics. The dataset includes 800 videos and 3,280 question-answer pairs, featuring tasks related to abstract concepts, symbolic elements, and multimodal integration, with varying levels of difficulty. Experimental results show that even state-of-the-art (SOTA) models, such as GPT-4o, achieve an average performance of 48.8% on tasks involving abstract concepts. Additionally, performance drops by 15% as task complexity increases, highlighting the challenges LVLMs face in maintaining consistent performance. Through this work, we hope to show the limitations of current LVLMs and offer insights into how they can more effectively emulate human cognitive processes in the future.

视频理解认知评估可控数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。