用文字描述代替视频分析,轻量高效地识别足球关键动作
Do We Need Large VLMs for Spotting Soccer Actions?
- 用三个专业分工的LLM替代复杂视觉模型,从解说文本中识别动作
- 零视频计算开销下达到接近顶尖视频模型的检测效果
- 适合追求低延迟、低成本的体育赛事实时分析场景
传统基于视频的足球动作识别高度依赖视觉输入,通常需要复杂且计算成本高昂的模型处理密集视频数据。本文提出一种从视频中心转向文本中心的新范式,利用大语言模型(LLMs)替代视觉-语言模型(VLMs),实现轻量化与可扩展性。我们假设专家解说中包含丰富的描述和上下文线索,足以可靠地识别比赛中的关键动作。为此,我们设计了一个由三个专业化LLM组成的系统,分别负责判别结果、激动程度和战术分析,共同完成动作定位任务。实验表明,该语言中心方法在检测关键比赛事件方面表现良好,性能接近当前最先进的视频基模型,同时实现了零视频处理计算开销,并在处理整场比赛时耗时相当。
原文摘要 · Abstract (English)
Traditional video-based tasks like soccer action spotting rely heavily on visual inputs, often requiring complex and computationally expensive models to process dense video data. We propose a shift from this video-centric approach to a text-based task, making it lightweight and scalable by utilizing Large Language Models (LLMs) instead of Vision-Language Models (VLMs). We posit that expert commentary, which provides rich descriptions and contextual cues contains sufficient information to reliably spot key actions in a match. To demonstrate this, we employ a system of three LLMs acting as judges specializing in outcome, excitement, and tactics for spotting actions in soccer matches. Our experiments show that this language-centric approach performs effectively in detecting critical match events coming close to state-of-the-art video-based spotters while using zero video processing compute and similar amount of time to process the entire match.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。