用视觉语言模型分析霹雳舞视频分类,发现编码器模型更优。
Breakdance Video classification in the age of Generative AI
- 采用视频基础模型编码器与解码器进行霹雳舞分类
- 编码器模型在预测任务中优于当前最先进水平
- 为舞蹈体育场景提供可复用的模型选择与调优方案
大型视觉语言模型近年来在多个体育场景中得到广泛应用,但多数研究集中于足球、板球、篮球等热门项目,主要聚焦于视觉问答、精彩片段生成等生成任务。本文探讨了现代视频基础模型(包括编码器和解码器)在小众但极受欢迎的舞蹈体育项目——霹雳舞中的适用性。结果表明,视频编码器模型在预测任务中仍显著优于当前最先进的视频语言模型。本文还提供了编码器模型选型建议,并对微调后的解码器模型在霹雳舞视频分类中的工作机制进行了深入分析。
原文摘要 · Abstract (English)
Large Vision Language models have seen huge application in several sports use-cases recently. Most of these works have been targeted towards a limited subset of popular sports like soccer, cricket, basketball etc; focusing on generative tasks like visual question answering, highlight generation. This work analyzes the applicability of the modern video foundation models (both encoder and decoder) for a very niche but hugely popular dance sports - breakdance. Our results show that Video Encoder models continue to outperform state-of-the-art Video Language Models for prediction tasks. We provide insights on how to choose the encoder model and provide a thorough analysis into the workings of a finetuned decoder model for breakdance video classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。