arXiv:2505.20124cs.CVcs.MM2025-05ACL被引 11

构建细粒度视频时序理解评测基准,全面评估动态视频理解能力。

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

  • 设计双任务评测框架:字幕生成与问答,覆盖多维度时序信息。
  • 揭示模型在动作描述、多主体理解和镜头运动感知上的显著不足。
  • 适合关注视频理解模型短板与评测方法的研究者使用。

视频独特地融合了时空元素,包括相机运动、场景、动作和属性及其随时间的动态关系。然而,现有视频理解基准常将这些属性分开处理或仅聚焦特定方面,忽视了视频内容的整体性。为此,我们提出 TUNA,一个面向密集动态视频的时序导向细粒度理解评测基准,包含互补的两个任务:字幕生成与问答。TUNA 涵盖多样视频场景与动态变化,并采用可解释且稳健的评估标准。我们在该基准上评估多个领先模型,实现跨多个维度的细粒度性能分析。结果揭示了当前模型在动作描述、多主体理解及相机运动敏感性方面的关键挑战,为改进视频理解模型提供了重要洞见。数据与代码已公开于 https://friedrichor.github.io/projects/TUNA。

原文摘要 · Abstract (English)

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often treat these properties separately or narrowly focus on specific aspects, overlooking the holistic nature of video content. To address this, we introduce TUNA, a temporal-oriented benchmark for fine-grained understanding on dense dynamic videos, with two complementary tasks: captioning and QA. Our TUNA features diverse video scenarios and dynamics, assisted by interpretable and robust evaluation criteria. We evaluate several leading models on our benchmark, providing fine-grained performance assessments across various dimensions. This evaluation reveals key challenges in video temporal understanding, such as limited action description, inadequate multi-subject understanding, and insensitivity to camera motion, offering valuable insights for improving video understanding models. The data and code are available at https://friedrichor.github.io/projects/TUNA.

视频理解时序评估细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。