arXiv:2608.19646cs.CV2026-08

首个基于进攻回合的篮球视频数据集,支持复杂视觉理解任务。

PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks

论文配图:PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks
图 1 · 摘自论文原文
  • 以完整进攻回合为样本构建数据集,保持事件时间连续性。
  • 包含1.1万段有效进攻片段和3.1万条标注事件,覆盖多种任务需求。
  • 适合研究战术分析、动作预测等复杂体育视觉理解问题。

近年来,体育场景中的视觉理解成为计算机视觉的热点。现有篮球视频数据集多以单一动作或活动为样本,既无法保留比赛事件的时间连续性,也无法支持动作预测等复杂任务。为此,本文构建了首个基于进攻回合的篮球视频数据集PL-NBA,每个样本包含一个完整的NBA进攻回合。数据集来自60场NBA比赛,包含11,000段有效的进攻回合视频片段和31,567条标注事件,涵盖球员姓名、描述文本、事件类型及时间戳。每段视频包含多个事件,保持事件连贯性,有助于战术分析。在事件识别、视频字幕生成、时间动作定位和动作预测等多个任务上进行实验,结果表明现有方法性能有限,验证了PL-NBA作为体育视频理解挑战性基准的有效性。

原文摘要 · Abstract (English)

Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue, this paper constructs the first possession-level basketball video dataset (PL-NBA), in which each sample is composed of a complete NBA offensive possession. Collected from 60 NBA games, PL-NBA contains 11,000 valid offensive possession clips and 31,567 annotated events with player names, captions, event types and timestamps. Each video clip includes multiple events and preserves the continuity of events, which is helpful for analysis of tactic. Experiment is conducted on multiple visual understanding tasks, including event recognition, video captioning, temporal action localization and action anticipation. Experimental results show that existing methods achieve limited performance on above four tasks, demonstrating that PL-NBA is a challenging benchmark for sports video understanding.

篮球分析视频理解数据集动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。