让网球视频理解从看动作到懂战术,证据链精准定位关键击球
TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

- 基于击球帧的证据链,联合推理战术与关键动作
- 在1.1万段回合中实现多任务精准预测,准确率超基线23%
- 适合研究体育智能分析、多模态大模型的从业者
体育视频理解正从事件识别转向解释动作如何共同影响比赛进程。现有方法或仅感知单个击球而忽略战术关联,或生成高层分析却缺乏底层事件支撑。为此,我们提出「击球证据锚定的战术推理」新任务,要求模型联合预测开放式答案、分层战术标签、支持性击球序列及决定性关键动作,并将每条证据击球精确对齐至球拍触球帧。我们进一步构建了TRACE基准,包含11,189段回合视频、41,485个击球事件、25,429个战术单元和11,189个问答对,统一整合细粒度击球属性、跨击球战术关系、分层战术标注及证据锚定的问题。基于此,我们提出TennisVAR(Tennis Video Action-chain Reasoner),一个遵循“事件-关系-证据-战术”推理范式的多模态大模型:事件解析模块将连续回合转化为显式击球序列;战术图引导的时间推理器联合建模比赛进展与同球员决策依赖,精准识别问题相关证据与关键动作。
原文摘要 · Abstract (English)
Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlying events. To bridge this perception-to-understanding gap, we formulate stroke-evidence-grounded tactical reasoning, a new rally-level task that requires models to jointly predict an open-ended answer, a hierarchical tactic label, an ordered sequence of supporting strokes, and decisive key actions, with each evidence stroke anchored to its racket-ball contact frame. We further introduce TRACE (Tactical Reasoning with Action-Chain Evidence in Tennis), a large-scale expert-annotated benchmark containing 11,189 rally videos, 41,485 stroke events, 25,429 tactical units, and 11,189 question-answer pairs, which unifies fine-grained stroke attributes, cross-stroke tactical relations, hierarchical tactic annotations, and evidence-grounded questions across factual perception, tactical understanding, and decision reasoning. Building on TRACE, we propose TennisVAR (Tennis Video Action-chain Reasoner), an evidence-grounded multimodal large language model that follows an "event-relation-evidence-tactic" reasoning paradigm, where an Event Parsing Module converts continuous rallies into explicit stroke-event sequences while a Tactical Graph-Guided Temporal Reasoner jointly models rally progression and same-player decision dependencies to identify question-relevant evidence and decisive actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。