让音乐表示与舞蹈动作对齐,提升节奏识别与生成能力
MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding
- 用身体运动信号训练音乐表示,通过对比学习捕捉精细节奏
- 在舞蹈生成任务中超越现有模型,且可迁移至节拍追踪等多任务
- 适合音乐-舞蹈生成、动作分析与跨模态检索的研究者
音乐既是听觉体验,也是身体感受,常通过舞蹈自然表达。然而现有音频表示普遍忽略身体维度,难以捕捉驱动动作的节奏与结构线索。本文提出MotionBeat,一种运动对齐的音乐表征学习框架。其引入两项新目标:具时序感知与节拍抖动负样本的体感对比损失(ECL),实现细粒度节奏区分;以及结构节奏对齐损失(SRAL),确保音乐重音与对应动作事件一致。架构上,采用周期性相位旋转建模循环节奏模式,并设计触点引导注意力机制突出与音乐重音同步的动作事件。实验表明,MotionBeat在音乐到舞蹈生成任务中优于现有先进音频编码器,并有效迁移至节拍追踪、音乐标签、流派与乐器分类、情绪识别及音视频检索等任务。
原文摘要 · Abstract (English)
Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their ability to capture rhythmic and structural cues that drive movement. We propose MotionBeat, a framework for motion-aligned music representation learning. MotionBeat is trained with two newly proposed objectives: the Embodied Contrastive Loss (ECL), an enhanced InfoNCE formulation with tempo-aware and beat-jitter negatives to achieve fine-grained rhythmic discrimination, and the Structural Rhythm Alignment Loss (SRAL), which ensures rhythm consistency by aligning music accents with corresponding motion events. Architecturally, MotionBeat introduces bar-equivariant phase rotations to capture cyclic rhythmic patterns and contact-guided attention to emphasize motion events synchronized with musical accents. Experiments show that MotionBeat outperforms state-of-the-art audio encoders in music-to-dance generation and transfers effectively to beat tracking, music tagging, genre and instrument classification, emotion recognition, and audio-visual retrieval. Our project demo page: https://motionbeat2025.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。