arXiv:2410.10818cs.CVcs.AI2024-10被引 90

新基准TemporalBench评估视频模型细粒度时间理解能力,揭示当前AI与人类差距达30%。

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

  • 基于2000段高质量人工标注,构建1万条视频问答对。
  • 顶尖模型GPT-4o在该基准上准确率仅38.5%,较人类水平低约30%。
  • 提出多二分类准确率(MBA)以纠正选择题中的误导性线索。

细粒度时间动态理解对多模态视频理解与生成至关重要。由于缺乏细粒度时间标注,现有视频基准大多类同静态图像基准,难以评估模型的时间理解能力。本文提出TemporalBench,一个专注于评估视频细粒度时间理解的新基准。该基准包含约10,000个视频问答对,源自约2,000段高质量人工标注,详尽描述视频片段中的时间动态。因此,该基准为评估动作频率、运动幅度、事件顺序等时间理解与推理能力提供了独特测试平台。同时支持视频问答、字幕生成、短/长视频理解等多种任务,以及多模态视频嵌入模型与文本生成模型的评估。结果表明,当前最优模型GPT-4o在TemporalBench上的问答准确率仅为38.5%,显示人工智能与人类在时间理解上存在约30%的显著差距。此外,我们发现多选题中大语言模型能通过负向描述的细微变化识别中心化提示,为此提出多重二分类准确率(MBA)以校正此类偏差。我们希望TemporalBench能推动提升模型时间推理能力的研究。数据集与评估代码将公开。

原文摘要 · Abstract (English)

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are incompetent at evaluating models for temporal understanding. In this paper, we introduce TemporalBench, a new benchmark dedicated to evaluating fine-grained temporal understanding in videos. TemporalBench consists of ~10K video question-answer pairs, derived from ~2K high-quality human annotations detailing the temporal dynamics in video clips. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as action frequency, motion magnitude, event order, etc. Moreover, it enables evaluations on various tasks like both video question answering and captioning, both short and long video understanding, as well as different models such as multimodal video embedding models and text generation models. Results show that state-of-the-art models like GPT-4o achieve only 38.5% question answering accuracy on TemporalBench, demonstrating a significant gap (~30%) between humans and AI in temporal understanding. Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a centralized description as a cue for its prediction, where we propose Multiple Binary Accuracy (MBA) to correct such bias. We hope that TemporalBench can foster research on improving models' temporal reasoning capabilities. Both dataset and evaluation code will be made available.

视频理解时间推理多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。