用多模态大模型分析网球回合动作,提升赛事策略洞察力
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
- 结合视觉、文本与音频输入,让模型理解网球回合动作序列
- 评估多模态大模型在识别单个动作及连续动作上的表现
- 探索融合传统模型与训练方法,提升分析准确率
近年来,大型语言模型(LLMs)的发展推动了多模态大模型(MLLMs)的兴起,使模型能够同时处理图像、视频和音频等多源信息。本研究旨在评估MLLMs在体育视频分析中的有效性,重点关注网球视频。尽管已有大量网球分析研究,但尚缺乏能准确理解并识别网球回合中动作序列的模型,这限制了其在更广泛运动分析领域的应用。因此,本研究重点评估MLLMs在分类网球动作以及识别回合中连续动作序列方面的能力。此外,还探索了通过不同训练策略及与传统模型结合等方式提升模型性能的可能性。
原文摘要 · Abstract (English)
The use of Large Language Models (LLMs) in recent years has also given rise to the development of Multimodal LLMs (MLLMs). These new MLLMs allow us to process images, videos and even audio alongside textual inputs. In this project, we aim to assess the effectiveness of MLLMs in analysing sports videos, focusing mainly on tennis videos. Despite research done on tennis analysis, there remains a gap in models that are able to understand and identify the sequence of events in a tennis rally, which would be useful in other fields of sports analytics. As such, we will mainly assess the MLLMs on their ability to fill this gap - to classify tennis actions, as well as their ability to identify these actions in a sequence of tennis actions in a rally. We further looked into ways we can improve the MLLMs' performance, including different training methods and even using them together with other traditional models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。