用离散动作符号实现舞蹈精准检索,支持大规模高效匹配。
Learning Quantised Structure-Preserving Motion Representations for Dance Fingerprinting
- 将人体姿态量化为结构化动作词汇,生成紧凑离散签名
- 在多舞种数据上实现跨风格高精度检索,支持未见编舞泛化
- 适合需要快速匹配与分析舞蹈动作的研究者和开发者
我们提出DANCEMATCH,一个端到端的动作驱动舞蹈检索框架,即从原始视频中直接识别语义相似的舞蹈编排(定义为舞蹈指纹)。现有方法依赖连续嵌入表示,难以索引、解释或扩展。DANCEMATCH通过骨架动作量化(SMQ)与时空变换器(STT)结合苹果CoMotion提取的姿态,构建捕捉舞蹈时空结构的紧凑离散签名,实现高效的大规模检索。系统还设计了基于直方图索引的子线性检索引擎(DRE),配合重排序实现精准匹配。为支持可复现研究,我们发布DANCETYPESBENCHMARK,一个带有量化动作标记的姿态对齐数据集。实验表明,该方法在多种舞蹈风格下表现稳健,并能有效泛化至未见编舞,为可扩展的动作指纹识别与定量编舞分析奠定基础。
原文摘要 · Abstract (English)
We present DANCEMATCH, an end-to-end framework for motion-based dance retrieval, the task of identifying semantically similar choreographies directly from raw video, defined as DANCE FINGERPRINTING. While existing motion analysis and retrieval methods can compare pose sequences, they rely on continuous embeddings that are difficult to index, interpret, or scale. In contrast, DANCEMATCH constructs compact, discrete motion signatures that capture the spatio-temporal structure of dance while enabling efficient large-scale retrieval. Our system integrates Skeleton Motion Quantisation (SMQ) with Spatio-Temporal Transformers (STT) to encode human poses, extracted via Apple CoMotion, into a structured motion vocabulary. We further design DANCE RETRIEVAL ENGINE (DRE), which performs sub-linear retrieval using a histogram-based index followed by re-ranking for refined matching. To facilitate reproducible research, we release DANCETYPESBENCHMARK, a pose-aligned dataset annotated with quantised motion tokens. Experiments demonstrate robust retrieval across diverse dance styles and strong generalisation to unseen choreographies, establishing a foundation for scalable motion fingerprinting and quantitative choreographic analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。