arXiv:2410.07752cs.CV2024-10被引 26

现有视频模型评测基准多依赖静态信息,新基准TVBench要求真正的时间推理能力。

Lost in Time: A New Temporal Benchmark for VideoLLMs

  • 设计新基准TVBench,强制模型理解视频时序动态变化。
  • 多数顶尖视频大模型在新基准上表现仅略高于随机猜测。
  • 适合评估真正具备时序理解能力的视频语言模型开发者使用。

大型语言模型与视觉模型结合后在视频理解任务中表现优异。然而,当前主流视频-语言评测基准存在三大问题:(i)单帧静态信息足以解题;(ii)问题与选项文本过于提示,无需视觉输入即可作答;(iii)多数问题仅靠常识即可回答,测试的是知识复现而非视频推理。此外,开放式问答评测依赖大模型自动评分,可靠性差。为此,本文提出TVBench——一个开源的视频多选题问答基准,经大量实验验证其需高水平时序理解能力。令人惊讶的是,大多数最新先进视频语言模型在该基准上表现接近随机水平,仅有少数如Qwen2-VL和Tarsier显著超越基线。

原文摘要 · Abstract (English)

Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating video models presents its own unique challenges, for which several benchmarks have been proposed. In this paper, we show that the currently most used video-language benchmarks can be solved without requiring much temporal reasoning. We identified three main issues in existing datasets: (i) static information from single frames is often sufficient to solve the tasks (ii) the text of the questions and candidate answers is overly informative, allowing models to answer correctly without relying on any visual input (iii) world knowledge alone can answer many of the questions, making the benchmarks a test of knowledge replication rather than video reasoning. In addition, we found that open-ended question-answering benchmarks for video understanding suffer from similar issues while the automatic evaluation process with LLMs is unreliable, making it an unsuitable alternative. As a solution, we propose TVBench, a novel open-source video multiple-choice question-answering benchmark, and demonstrate through extensive evaluations that it requires a high level of temporal understanding. Surprisingly, we find that most recent state-of-the-art video-language models perform similarly to random performance on TVBench, with only a few models such as Qwen2-VL, and Tarsier clearly surpassing this baseline.

视频理解评测基准时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。