构建多模态篮球理解基准,提升模型综合推理能力
Towards Comprehensive Basketball Understanding

- 设计可组合的领域专用工具链,分步完成篮球认知任务
- 在7980个跨模态问题上测试,现有大模型整合能力弱
- 适合研究体育视频理解、多智能体协同系统的学者
理解一场篮球比赛需要识别事件、定位动作、辨识球员,并将其与结构化比赛知识关联。现有评测基准多单独评估某项能力,忽视能力间的交互。我们提出BasketballBench,一个包含7,980个问题的多模态基准,涵盖文本、图像、视频十类任务,数据源自2025-2026 NBA赛季,包含官方逐回合记录、530名现役球员的名单与档案,以及2,501段进攻回合级转播片段。我们进一步提出BasketballSkills,一种由八个篮球专用感知与检索工具组成的智能体,基于四种可复用技能定义工具顺序、证据绑定和停止条件。实验表明,当前多模态大模型在需多能力融合的问题上表现不佳,而BasketballSkills显著优于它们,验证了显式组合领域专用能力对全面篮球理解的有效性。
原文摘要 · Abstract (English)
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。