Poze用少量数据实现专业级运动技术反馈,让普通人也能获得教练级指导。
Poze: Sports Technique Feedback under Data Constraints
- 结合姿态估计与序列比对,低数据下模拟教练分析动作
- 在视频问答任务中准确率比GPT4V高70%,比LLaVAv1.6 7b高196%
- 适合缺乏专业教练资源的运动爱好者和训练场景
专业教练对运动技术发展至关重要,但经济门槛常使许多人难以获得。为弥合这一差距,我们提出Poze,一种创新的视频处理框架,可提供人体运动反馈,模拟专业教练的洞察。Poze融合姿态估计与序列对比,在数据有限条件下仍表现优异。在视频问答框架中,其准确率分别较GPT4V提升70%,较LLaVAv1.6 7b提升196%。
原文摘要 · Abstract (English)
Access to expert coaching is essential for developing technique in sports, yet economic barriers often place it out of reach for many enthusiasts. To bridge this gap, we introduce Poze, an innovative video processing framework that provides feedback on human motion, emulating the insights of a professional coach. Poze combines pose estimation with sequence comparison and is optimized to function effectively with minimal data. Poze surpasses state-of-the-art vision-language models in video question-answering frameworks, achieving 70% and 196% increase in accuracy over GPT4V and LLaVAv1.6 7b, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。