构建首个实时篮球解说大规模基准,支持细粒度事件追踪与连贯描述。
NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams

- 提出因果双阶段框架,先补全事件再定位,结合球为中心的语义对齐
- 包含307.5小时直播数据与约3.5万条时间对齐事件标注,覆盖球员身份、动作链等细粒度信息
- 适用于实时体育视频理解与生成研究,推动在线解说系统发展
实时篮球解说需在事件充分显现前完成描述,但现有方法多针对预分割片段或完整视频,难以适应连续流。现有数据集对球员身份、细粒度动作、事件属性及连贯事件链的标注有限,制约了解说的事实丰富性。为此,我们推出NBA_Streaming,一个大规模在线细粒度篮球解说生成基准。该数据集包含307.5小时篮球直播和约3.5万条时间对齐事件,涵盖事件边界、球员身份、细粒度动作、事件链及自然语言解说。通过从孤立片段转向连续流,NBA_Streaming实现了对事件定位、响应可靠性、事实一致性及解说质量在因果约束下的统一评估。我们进一步提出一种因果双阶段框架,结合补全优先的定位与以球为中心的语义对齐,使系统能从观测流中识别完整事件,并整合场景、事件、身份与动作线索进行解说生成。大量实验表明,该任务难度显著,现有基线在在线时机、事实一致性和细粒度描述上表现不佳。我们的框架持续优于强基线,剩余差距凸显了NBA_Streaming作为流式体育视频理解与生成重要基准的价值。代码与数据将在论文录用后公开。
原文摘要 · Abstract (English)
Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. However, existing methods are primarily designed for pre-segmented clips or complete videos, making them unsuitable for continuous streams. Existing datasets also provide limited supervision for player identities, fine-grained actions, event attributes, and coherent event chains, restricting the factual richness of generated commentary. To address these limitations, we introduce NBA_Streaming, a large-scale benchmark for online fine-grained basketball commentary generation. It contains 307.5 hours of basketball broadcasts and approximately 35K temporally aligned events, with annotations of event boundaries, player identities, fine-grained actions, event chains, and natural-language commentary. By moving from isolated clips to continuous streams, NBA_Streaming enables unified evaluation of event localization, response reliability, factual grounding, and commentary quality under causal constraints. We further propose a causal two-stage framework that combines completion-first localization with ball-centric semantic grounding, enabling the system to identify complete events from observed streams and organize scene, event, identity, and action cues for commentary generation. Extensive experiments reveal the difficulty of NBA_Streaming, where existing baselines struggle with online timing, factual grounding, and fine-grained description. Our framework consistently improves over strong alternatives, while the remaining gap highlights NBA_Streaming as a valuable benchmark for streaming sports video understanding and generation. The code and data will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。