用视觉模型自动分析非洲舞步,识别动作和空间使用差异。
AfroBeats Dance Movement Analysis Using Computer Vision: A Proof-of-Concept Framework Combining YOLO and Segment Anything Model
- 结合YOLO与SAM模型,实现无标记舞蹈动作检测与像素级分割。
- 主舞者动作多23%,强度高37%,占空间多42%,量化差异明显。
- 适合对舞蹈量化、文化研究或计算机视觉应用感兴趣的读者。
本文提出一种基于现代计算机视觉技术的舞蹈动作自动化分析初步框架。通过整合YOLOv8和v11进行舞者检测,利用Segment Anything Model(SAM)实现精准分割,可在无特殊设备或标记的情况下,对视频中的舞者运动进行追踪与量化。该方法可识别舞者、统计离散舞步数量、计算空间覆盖模式,并测量表演序列中的节奏一致性。在一段49秒的加纳非洲舞视频上测试,系统在人工标注样本中达到约94%的检测精度和89%的召回率。SAM提供的像素级分割与人工检查相比,交并比约为83%,能捕捉超出边界框表示的身体姿态变化。初步案例分析显示,系统识别出的主舞者比次级舞者多执行23%的舞步,运动强度高出37%,使用的表演空间多42%。然而,本研究仍属早期探索,存在仅基于单个视频验证、缺乏系统性真值标注以及未与现有姿态估计方法对比等局限。本文旨在展示技术可行性,揭示定量舞蹈指标的潜力,并为未来系统性验证研究奠定基础。
原文摘要 · Abstract (English)
This paper presents a preliminary investigation into automated dance movement analysis using contemporary computer vision techniques. We propose a proof-of-concept framework that integrates YOLOv8 and v11 for dancer detection with the Segment Anything Model (SAM) for precise segmentation, enabling the tracking and quantification of dancer movements in video recordings without specialized equipment or markers. Our approach identifies dancers within video frames, counts discrete dance steps, calculates spatial coverage patterns, and measures rhythm consistency across performance sequences. Testing this framework on a single 49-second recording of Ghanaian AfroBeats dance demonstrates technical feasibility, with the system achieving approximately 94% detection precision and 89% recall on manually inspected samples. The pixel-level segmentation provided by SAM, achieving approximately 83% intersection-over-union with visual inspection, enables motion quantification that captures body configuration changes beyond what bounding-box approaches can represent. Analysis of this preliminary case study indicates that the dancer classified as primary by our system executed 23% more steps with 37% higher motion intensity and utilized 42% more performance space compared to dancers classified as secondary. However, this work represents an early-stage investigation with substantial limitations including single-video validation, absence of systematic ground truth annotations, and lack of comparison with existing pose estimation methods. We present this framework to demonstrate technical feasibility, identify promising directions for quantitative dance metrics, and establish a foundation for future systematic validation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。