用Transformer生成符合音乐结构的3D指挥动作,更精准同步且自然。
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation

- 基于Transformer建模音乐与手势的时序关系,自回归预测3D人体姿态。
- 在专业指挥动作数据集上,生成动作与音乐对齐度显著优于基线模型。
- 引入多模态嵌入评估,可量化音乐与手势的艺术匹配度,适合作曲家与编舞者使用。
从音乐生成富有表现力的指挥手势是一项具有挑战性的跨模态运动合成任务:输出需遵循长时序音乐结构,保持节拍级同步,并呈现为精细的3D人体动作。现有研究常受限于稀疏姿态表示、小规模数据或无法直接衡量音乐与手势互相对齐的评估协议。本文提出TransConductor,一种基于Transformer的音乐驱动指挥动作生成框架。我们构建了ConductorMotion数据构造流程,从指挥视频中恢复详细身体运动,形成面向专业指挥动作的数据集。给定音频提取的声学描述符和初始姿态,TransConductor通过时序音乐编码器与时序手势解码器,自回归预测SMPL姿态参数。为更好评估艺术对应性,我们进一步构建了基于检索的评估模型,将音乐与手势嵌入共享空间,提供FID、模态距离、多模态距离和多样性指标。实验表明,TransConductor在生成质量上优于舞蹈与指挥生成基线,消融实验证明Transformer主干与提出的对齐损失的有效性。
原文摘要 · Abstract (English)
Generating expressive conducting gestures from music is a challenging cross-modal motion synthesis problem: the output must follow long-range musical structure, preserve beat-level synchronization, and remain plausible as a fine-grained 3D human performance. Existing conducting-motion studies are often limited by sparse pose representations, small-scale data, or evaluation protocols that do not directly measure whether music and gesture are mutually aligned. This paper presents TransConductor, a Transformer-based framework for music-driven conducting gesture generation. We introduce ConductorMotion, a SMPL-parameter data construction pipeline that recovers detailed body motion from conducting videos and forms a dataset targeted at professional conducting gestures. Given acoustic descriptors extracted from audio and an initial pose, TransConductor uses a Trans-Temporal Music Encoder and a Trans-Temporal Conducting Gesture Decoder to autoregressively predict SMPL pose parameters. To better assess artistic correspondence, we further build a retrieval-based evaluation model that embeds music and gestures into a shared space and yields FID, modality distance, multi-modality distance, and diversity metrics. Experiments show that TransConductor outperforms dance-generation and conducting-generation baselines, while ablations verify the benefits of the Transformer backbone and the proposed alignment loss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。