构建首个体育赛事视频解说基准,评估模型细粒度视觉理解能力。
SCBench: A Sports Commentary Benchmark for Video LLMs
- 设计六维评分体系与GPT评估法,精准衡量解说生成质量。
- 发布包含5775段标注视频的CommentarySet数据集,覆盖复杂视觉场景。
- 发现InternVL-Chat-2在5.44分上表现最佳,显著优于其他模型。
近年来,视频大语言模型(Video LLMs)在学术界和工业界取得显著进展,但对其性能的评估方法仍十分有限,尤其缺乏对细粒度时序视觉能力的评测。现有基准多采用简单视频(如带字幕的电影片段),模型仅需处理少数帧即可理解内容;同时任务形式单一,仅限问答或选择题,无法检验模型生成深度、精确文本的能力。体育视频具有复杂的视觉信息、连续事件和情感化表达,是检验视频理解能力的理想任务。为此,我们提出体育视频解说生成新任务,并构建了SCBench基准。该基准包含:(1) 专为本任务设计的六维评价指标SCORES,结合GPT评估方法;(2) CommentarySet数据集,含5,775个标注视频片段及真实标注。基于SCBench,我们对多个Video LLM(如VILA、Video-LLaVA)及思维链基线进行了全面评估。结果表明,InternVL-Chat-2以5.44分表现最优,领先第二名1.04分。本工作为未来复杂视觉理解研究提供了新视角,数据集将尽快开源。
原文摘要 · Abstract (English)
Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance of different Video LLMs, especially their fine-grained, temporal visual capabilities, remain very limited. On one hand, current benchmarks use relatively simple videos (e.g., subtitled movie clips) where the model can understand the entire video by processing just a few frames. On the other hand, their datasets lack diversity in task format, comprising only QA or multi-choice QA, which overlooks the models' capacity for generating in-depth and precise texts. Sports videos, which feature intricate visual information, sequential events, and emotionally charged commentary, present a critical challenge for Video LLMs, making sports commentary an ideal benchmarking task. Inspired by these challenges, we propose a novel task: sports video commentary generation, developed $\textbf{SCBench}$ for Video LLMs. To construct such a benchmark, we introduce (1) $\textbf{SCORES}$, a six-dimensional metric specifically designed for our task, upon which we propose a GPT-based evaluation method, and (2) $\textbf{CommentarySet}$, a dataset consisting of 5,775 annotated video clips and ground-truth labels tailored to our metric. Based on SCBench, we conduct comprehensive evaluations on multiple Video LLMs (e.g. VILA, Video-LLaVA, etc.) and chain-of-thought baseline methods. Our results found that InternVL-Chat-2 achieves the best performance with 5.44, surpassing the second-best by 1.04. Our work provides a fresh perspective for future research, aiming to enhance models' overall capabilities in complex visual understanding tasks. Our dataset will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。