arXiv:2410.23266cs.CVcs.AI2024-10被引 74

新基准TOMATO揭示大模型视频时序推理能力被高估

TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

  • 设计三原则评估多帧依赖、帧序敏感性与信息差异
  • 人类与最佳模型在时序推理上差距达57.3%
  • 适合关注视频理解真实能力的开发者与研究者

现有基准常显示多模态基础模型在利用时间上下文进行视频理解方面表现优异,但其真正的视觉时序推理能力究竟如何?我们研究发现,当前任务常可通过单帧、少数帧或无序帧解决,导致模型能力被高估。为此,我们提出三个评估原则及对应指标:(1)多帧增益,(2)帧序敏感性,(3)帧信息差异。基于此,构建TOMATO(Temporal Reasoning Multimodal Evaluation)新基准,包含1,484个人工标注的问题,覆盖6类任务(动作计数、方向、旋转、形状与趋势、速度与频率、视觉线索),应用于1,417个视频(含805个自录/生成视频),涵盖以人为中心、真实世界与模拟场景。全面评估显示,最佳模型与人类表现差距达57.3%。深入分析揭示,尽管模型可准确识别孤立帧事件,却无法将帧序列视为连续动态过程。TOMATO将成为下一代多模态模型的重要评测平台,并呼吁社区发展能理解人类世界动态的视频智能系统。

原文摘要 · Abstract (English)

Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understanding. However, how well do the models truly perform visual temporal reasoning? Our study of existing benchmarks shows that this capability of MFMs is likely overestimated as many questions can be solved by using a single, few, or out-of-order frames. To systematically examine current visual temporal reasoning tasks, we propose three principles with corresponding metrics: (1) Multi-Frame Gain, (2) Frame Order Sensitivity, and (3) Frame Information Disparity. Following these principles, we introduce TOMATO, Temporal Reasoning Multimodal Evaluation, a novel benchmark crafted to rigorously assess MFMs' temporal reasoning capabilities in video understanding. TOMATO comprises 1,484 carefully curated, human-annotated questions spanning six tasks (i.e., action count, direction, rotation, shape & trend, velocity & frequency, and visual cues), applied to 1,417 videos, including 805 self-recorded and -generated videos, that encompass human-centric, real-world, and simulated scenarios. Our comprehensive evaluation reveals a human-model performance gap of 57.3% with the best-performing model. Moreover, our in-depth analysis uncovers more fundamental limitations beyond this gap in current MFMs. While they can accurately recognize events in isolated frames, they fail to interpret these frames as a continuous sequence. We believe TOMATO will serve as a crucial testbed for evaluating the next-generation MFMs and as a call to the community to develop AI systems capable of comprehending human world dynamics through the video modality.

视频理解时序推理多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。