首个面向体育场景的视觉语言模型空间智能评测基准。
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
- 构建包含百万级问答对的CourtSI数据集,覆盖计数、距离、定位等空间推理任务。
- 在3686个经人工验证的题目上测试25个模型,发现人类与AI仍有显著差距。
- 微调后模型在新运动场景中表现提升,适用于体育解说生成等应用。
体育赛事长期吸引广泛关注,因其展现了人类体能与认知能力的极限。随着视觉语言模型(VLMs)对空间智能关注度提升,体育场景成为理解高强度人体运动与动态物体交互的理想测试平台。为此,我们提出CourtSI——首个面向体育场景的大规模空间智能数据集。CourtSI包含超过100万条问答对,基于全面分类体系,系统涵盖空间计数、距离测量、定位及关系推理,覆盖羽毛球、网球、乒乓球等典型隔网运动。利用明确的场地几何结构作为度量锚点,我们开发半自动数据引擎重建体育场景,实现CourtSI的可扩展构建。此外,我们引入CourtSI-Bench,一个包含3,686个经严格人工验证的问答对的高质量评估基准。我们在CourtSI-Bench上评估了25个商用与开源VLMs,揭示了当前模型与人类之间的性能差距,以及现有基准在泛化能力上的局限性。这些发现表明,体育场景暴露了现有基准未能充分捕捉的空间智能缺陷。进一步地,将Qwen3-VL-8B在CourtSI上微调后,在CourtSI-Bench上的准确率提升23.5个百分点;该优化模型在针对相似但未见过的运动构建的CourtSI-Ext上也表现出良好泛化能力,并增强了空间感知的解说生成效果。这些成果表明,CourtSI为推动VLM在体育领域中的空间智能发展提供了可扩展路径。
原文摘要 · Abstract (English)
Sports have long attracted broad attention as they push the limits of human physical and cognitive capabilities. Amid growing interest in spatial intelligence for vision-language models (VLMs), sports provide a natural testbed for understanding high-intensity human motion and dynamic object interactions. To this end, we present CourtSI, the first large-scale spatial intelligence dataset tailored to sports scenarios. CourtSI contains over 1M QA pairs, organized under a holistic taxonomy that systematically covers spatial counting, distance measurement, localization, and relational reasoning, across representative net sports including badminton, tennis, and table tennis. Leveraging well-defined court geometry as metric anchors, we develop a semi-automatic data engine to reconstruct sports scenes, enabling scalable curation of CourtSI. In addition, we introduce CourtSI-Bench, a high-quality evaluation benchmark comprising 3,686 QA pairs with rigorous human verification. We evaluate 25 proprietary and open-source VLMs on CourtSI-Bench, revealing a remaining human-AI performance gap and limited generalization from existing spatial intelligence benchmarks. These findings indicate that sports scenarios expose limitations in spatial intelligence capabilities captured by existing benchmarks. Further, fine-tuning Qwen3-VL-8B on CourtSI improves accuracy on CourtSI-Bench by 23.5 percentage points. The adapted model also generalizes effectively to CourtSI-Ext, an evaluation set built on a similar but unseen sport, and demonstrates enhanced spatial-aware commentary generation. Together, these findings demonstrate that CourtSI provides a scalable pathway toward advancing spatial intelligence of VLMs in sports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。