arXiv:2511.06499cs.CV2025-11被引 14

构建首个多运动大模型推理基准,评估视觉理解与规则推理能力。

SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

  • 设计分层问答体系,从判罚识别到战术解释逐级挑战模型推理
  • 包含4789张图像、2052段视频及6841条思维链标注,支持精细评估
  • 适合研究多模态推理、体育智能或视觉接地的学者使用

深入理解体育运动需要精细的视觉感知与基于规则的推理相结合——这对当前多模态模型构成巨大挑战。模型必须掌握三项核心能力:捕捉细微视觉细节、应用抽象运动规则知识,并将知识准确关联到具体视觉证据。现有体育评测数据集要么仅覆盖单一运动,要么缺乏细致的推理链条与精确的视觉定位,难以在多运动场景下有效评估这些关键能力。为此,我们提出SportR,首个大规模多运动多模态大模型推理基准,用于训练与评估模型在体育智能中的基础推理能力。该基准包含4,789张图像和2,052段视频,通过渐进式问题-答案对结构,探测从简单违规识别到复杂罚则预测等多层次推理能力。针对需多步推理的任务(如判定罚则或解释战术),我们提供6,841条高质量人工编写的思维链(Chain of Thought)标注。同时,基准涵盖图像与视频双模态,并提供手动边界框标注以直接测试图像部分的视觉定位能力。大量实验表明,该基准极具挑战性:主流基线模型在最困难任务上表现不佳。尽管通过监督微调和强化学习在该数据上训练可提升性能,但得分仍相对较低,凸显当前模型能力的巨大差距。SportR为社区提供了新的挑战,是推动多模态体育推理研究的关键资源。

原文摘要 · Abstract (English)

Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical capabilities: perceiving nuanced visual details, applying abstract sport rule knowledge, and grounding that knowledge in specific visual evidence. Current sports benchmarks either cover single sports or lack the detailed reasoning chains and precise visual grounding needed to robustly evaluate these core capabilities in a multi-sport context. To address this gap, we introduce SportR, the first multi-sports large-scale benchmark designed to train and evaluate MLLMs on the fundamental reasoning required for sports intelligence. Our benchmark provides a dataset of 4,789 images and 2,052 videos. To enable granular evaluation, we structure our benchmark around a progressive hierarchy of question-answer pairs designed to probe reasoning at increasing depths - from simple infraction identification to complex penalty prediction. For the most advanced tasks requiring multi-step reasoning, such as determining penalties or explaining tactics, we provide 6,841 high-quality, human-authored Chain of Thought annotations. In addition, our benchmark incorporates both image and video modalities and provides manual bounding box annotations to test visual grounding in the image part directly. Extensive experiments demonstrate the profound difficulty of our benchmark. State-of-the-art baseline models perform poorly on our most challenging tasks. While training on our data via Supervised Fine-Tuning and Reinforcement Learning improves these scores, they remain relatively low, highlighting a significant gap in current model capabilities. SportR presents a new challenge for the community, providing a critical resource to drive future research in multimodal sports reasoning.

多模态推理体育智能视觉接地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。