首个网球回合理解评测基准,检验多模态大模型在高速运动中的表现
TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
- 将每个回合拆解为连续击球事件序列,构建时序化评估框架
- 17个模型在2527个人工验证问题上测试,发现采样密度需任务适配
- 提升时间定位能力是增强推理的关键,适合体育视频分析研究者
多模态大语言模型(MLLMs)在通用视频理解中表现优异,但在网球等高频快速运动中仍面临挑战,因其回合片段短而信息密集。为系统评估此类场景下的MLLM表现,我们提出TennisTV——首个且最全面的网球视频理解基准。TennisTV将每个回合建模为连续击球事件的时间有序序列,采用自动化流程进行数据筛选与问题生成。涵盖从击球级到回合级共8项任务,包含2527个经人工验证的问题。对17个代表性MLLM进行评估,首次实现网球视频理解的系统性分析。结果揭示两个关键洞见:(i) 帧采样密度应根据任务需求动态调整与平衡;(ii) 提升时间定位能力对强化模型推理至关重要。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) excel at general video understanding but struggle with fast, high-frequency sports like tennis, where rally clips are short yet information-dense. To systematically evaluate MLLMs in this challenging domain, we present TennisTV, the first and most comprehensive benchmark for tennis video understanding. TennisTV models each rally as a temporal-ordered sequence of consecutive stroke events, using automated pipelines for filtering and question generation. It covers 8 tasks from the stroke level to the rally level and includes 2527 human-verified questions. Evaluating 17 representative MLLMs, we provide the first systematic assessment of tennis video understanding. Results yield two key insights: (i) frame-sampling density should be tailored and balanced across tasks, and (ii) improving temporal grounding is essential for stronger reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。