arXiv:2604.04419cs.CV2026-04

首个聚焦拳击的自动解说数据集,评测模型对战术分析与节奏把控能力。

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing

  • 构建拳击解说分类体系,区分动作描述、战术分析与背景信息。
  • 提出类型条件生成与解说节奏评估两项新评测,揭示模型短板。
  • 适合研究体育智能解说、多模态理解与细粒度事件感知的学者。

近期多模态大语言模型在通用视频理解方面表现强劲,推动了自动体育解说生成的研究兴趣。然而,现有基准仅关注足球、篮球等团体运动,完全忽视了格斗类运动。格斗类运动具有独特挑战:关键动作在毫秒级内完成,视觉差异细微但语义决定性强,且专业解说中战术分析占比远高于团队运动。本文提出BoxComm,一个大规模数据集,包含445场世界拳击锦标赛比赛视频及超过52,000条来自专业转播的解说文本。我们设计了一套结构化解说分类体系,将每条解说句标注为动作描述、战术分析或背景信息,这是首个支持类别级标注的体育解说基准。基于此分类体系,我们引入两项新颖且互补的评估方法:(1) 类型条件生成,检验模型能否根据视频上下文生成指定类型的准确解说;(2) 解说节奏评估,衡量自由生成解说在连续视频段落中是否具备恰当的时间节奏与类型分布,捕捉了此前基准未覆盖的解说能力维度。在多个先进MLLM上的实验表明,当前模型在两项评估上均表现不佳。我们进一步提出EIC-Gen基线,通过检测拳击动作提供结构化动作提示,带来稳定提升,凸显了对瞬时微小事件感知在格斗类运动解说中的重要性。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) have shown strong capabilities in general video understanding, driving growing interest in automatic sports commentary generation. However, existing benchmarks for this task focus exclusively on team sports such as soccer and basketball, leaving combat sports entirely unexplored. Notably, combat sports present distinct challenges: critical actions unfold within milliseconds with visually subtle yet semantically decisive differences, and professional commentary contains a substantially higher proportion of tactical analysis compared to team sports. In this paper, we present BoxComm, a large-scale dataset comprising 445 World Boxing Championship match videos with over 52K commentary sentences from professional broadcasts. We propose a structured commentary taxonomy that categorizes each sentence into play-by-play, tactical, or contextual, providing the first category-level annotation for sports commentary benchmarks. Building on this taxonomy, we introduce two novel and complementary evaluations tailored to sports commentary generation: (1) category-conditioned generation, which evaluates whether models can produce accurate commentary of a specified type given video context; and (2) commentary rhythm assessment, which measures whether freely generated commentary exhibits appropriate temporal pacing and type distribution over continuous video segments, capturing a dimension of commentary competence that prior benchmarks have not addressed. Experiments on multiple state-of-the-art MLLMs reveal that current models struggle on both evaluations. We further propose EIC-Gen, an improved baseline incorporating detected punch events to supply structured action cues, yielding consistent gains and highlighting the importance of perceiving fleeting and subtle events for combat sports commentary.

体育解说多模态拳击评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。