构建多模态大模型体育理解评测基准,覆盖从识球到判罚的多层次推理任务。
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
- 分文本与视频双模块,测试模型对规则和策略的理解能力
- 视频任务中硬题准确率最高仅52.6%,显示模型仍存巨大提升空间
- 适合评估多模态大模型在体育场景中的深度推理性能
多模态大语言模型(MLLMs)正通过融合文本与视觉信息提升对复杂体育场景的推理能力。为全面评估其表现,我们提出SPORTU,一个涵盖多层级体育推理任务的评测基准。该基准包含两个核心组件:SPORTU-text含900道多选题及人工标注解释,用于测试仅基于问答的规则理解与策略分析能力;SPORTU-video包含7种运动的1,701段慢动作视频和12,048个问答对,覆盖从简单识别到复杂判罚与规则应用的任务。我们在SPORTU-text上评估了四种主流大模型,采用少样本学习与思维链(CoT)提示,结果显示GPT-4o准确率达71%,仍低于人类水平。在SPORTU-video部分,共测试7个专有与6个开源多模态模型,结果表明模型在需要深度推理与规则理解的难题上表现不佳,其中最佳模型Claude-3.5-Sonnet在难题上的准确率仅为52.6%。该工作期望推动多模态模型在体育理解领域的评测与进步。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPORTU, a benchmark designed to assess MLLMs across multi-level sports reasoning tasks. SPORTU comprises two key components: SPORTU-text, featuring 900 multiple-choice questions with human-annotated explanations for rule comprehension and strategy understanding. This component focuses on testing models' ability to reason about sports solely through question-answering (QA), without requiring visual inputs; SPORTU-video, consisting of 1,701 slow-motion video clips across 7 different sports and 12,048 QA pairs, designed to assess multi-level reasoning, from simple sports recognition to complex tasks like foul detection and rule application. We evaluate four prevalent LLMs mainly utilizing few-shot learning paradigms supplemented by chain-of-thought (CoT) prompting on the SPORTU-text part. We evaluate four LLMs using few-shot learning and chain-of-thought (CoT) prompting on SPORTU-text. GPT-4o achieves the highest accuracy of 71%, but still falls short of human-level performance, highlighting room for improvement in rule comprehension and reasoning. The evaluation for the SPORTU-video part includes 7 proprietary and 6 open-source MLLMs. Experiments show that models fall short on hard tasks that require deep reasoning and rule-based understanding. Claude-3.5-Sonnet performs the best with only 52.6% accuracy on the hard task, showing large room for improvement. We hope that SPORTU will serve as a critical step toward evaluating models' capabilities in sports understanding and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。