用多个协作智能体分析羽毛球视频,实现从细节到全局的灵活理解
COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
- 设计多智能体系统,每个智能体专精不同分析任务
- 可动态组合成短时问答与长时摘要等不同分析流程
- 提升模型可解释性与跨任务泛化能力,适合体育视频研究者
智能体育视频分析需要对时间上下文有全面理解,涵盖微观动作到宏观战术。现有端到端模型常因难以处理时间层次结构,导致泛化能力差、新任务开发成本高且可解释性弱。为此,我们提出一种可重构的多智能体系统(MAS)作为体育视频理解的基础框架。每个智能体作为特定分析维度的“认知工具”,系统架构不局限于单一时间尺度或任务。通过迭代调用和灵活组合,该框架可构建适应性分析流水线,支持短期推理(如发球回合问答)与长期生成摘要(如比赛总结)。我们在羽毛球分析中验证其适用性,展示其在细粒度事件检测与全局语义组织间的桥梁作用。本工作推动了面向跨任务、可扩展、可解释的体育视频智能新范式。项目主页:https://aiden1020.github.io/COACH-project-page
原文摘要 · Abstract (English)
Intelligent sports video analysis demands a comprehensive understanding of temporal context, from micro-level actions to macro-level game strategies. Existing end-to-end models often struggle with this temporal hierarchy, offering solutions that lack generalization, incur high development costs for new tasks, and suffer from poor interpretability. To overcome these limitations, we propose a reconfigurable Multi-Agent System (MAS) as a foundational framework for sports video understanding. In our system, each agent functions as a distinct "cognitive tool" specializing in a specific aspect of analysis. The system's architecture is not confined to a single temporal dimension or task. By leveraging iterative invocation and flexible composition of these agents, our framework can construct adaptive pipelines for both short-term analytic reasoning (e.g., Rally QA) and long-term generative summarization (e.g., match summaries). We demonstrate the adaptability of this framework using two representative tasks in badminton analysis, showcasing its ability to bridge fine-grained event detection and global semantic organization. This work presents a paradigm shift towards a flexible, scalable, and interpretable system for robust, cross-task sports video intelligence. The project homepage is available at https://aiden1020.github.io/COACH-project-page
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。