arXiv:2504.09282cs.CV2025-04ICCV被引 9

首个面向广告视频的多模态模型评测数据集,挑战模型时序理解能力。

VideoAds for Fast-Paced Video Understanding

  • 构建首个专注广告视频的多模态评测数据集,含复杂时间结构与人工标注问题。
  • 开源模型Qwen2.5-VL-72B在该数据集上达73.35%准确率,超越GPT-4o和Gemini-1.5 Pro。
  • 人类专家准确率达94.27%,凸显当前模型在时序推理上的明显不足。

广告视频包含高质量视觉、文本和上下文线索,旨在吸引观众,其叙事结构严谨、场景切换迅速,比同长度普通视频更复杂,对多模态大语言模型(MLLM)构成重大挑战。本文提出VideoAds,首个专为评估MLLM在广告视频上表现而设计的数据集。该数据集包含精心筛选的广告视频,具有复杂的时序结构,并配有三类核心任务的人工标注问题:视觉定位、视频摘要和视觉推理。我们提出量化指标,对比VideoAds与现有基准的视频复杂度。大量实验表明,开源模型Qwen2.5-VL-72B在VideoAds上达到73.35%准确率,优于GPT-4o(66.82%)和Gemini-1.5 Pro(69.66%);两款闭源模型在视频摘要与推理任务中表现落后,但在视觉定位上最佳。值得注意的是,人类专家可达94.27%的准确率。结果表明,亟需提升MLLM的时序建模能力,且VideoAds有望成为未来高帧率视频理解研究的关键基准。数据集与评估代码将公开发布于https://videoadsbenchmark.netlify.app。

原文摘要 · Abstract (English)

Advertisement videos serve as a rich and valuable source of purpose-driven information, encompassing high-quality visual, textual, and contextual cues designed to engage viewers. They are often more complex than general videos of similar duration due to their structured narratives and rapid scene transitions, posing significant challenges to multi-modal large language models (MLLMs). In this work, we introduce VideoAds, the first dataset tailored for benchmarking the performance of MLLMs on advertisement videos. VideoAds comprises well-curated advertisement videos with complex temporal structures, accompanied by \textbf{manually} annotated diverse questions across three core tasks: visual finding, video summary, and visual reasoning. We propose a quantitative measure to compare VideoAds against existing benchmarks in terms of video complexity. Through extensive experiments, we find that Qwen2.5-VL-72B, an opensource MLLM, achieves 73.35\% accuracy on VideoAds, outperforming GPT-4o (66.82\%) and Gemini-1.5 Pro (69.66\%); the two proprietary models especially fall behind the opensource model in video summarization and reasoning, but perform the best in visual finding. Notably, human experts easily achieve a remarkable accuracy of 94.27\%. These results underscore the necessity of advancing MLLMs' temporal modeling capabilities and highlight VideoAds as a potentially pivotal benchmark for future research in understanding video that requires high FPS sampling. The dataset and evaluation code will be publicly available at https://videoadsbenchmark.netlify.app.

视频理解多模态广告视频时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。