arXiv:2604.04395cs.CVcs.MM2026-04被引 6

用新模型和数据集实现高精度3D指挥动作生成,支持高效长序列合成。

BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion

论文配图:BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
图 1 · 摘自论文原文
  • 融合BiMamba与Transformer的混合架构,兼顾长序列效率与跨模态对齐。
  • 在10小时的精细3D指挥数据集上达到当前最优生成质量。
  • 支持无需训练的动作编辑,适合人机协作与数字人动画应用。

3D指挥动作生成旨在从音乐中合成精细的指挥动作,广泛应用于音乐教育、虚拟演出、数字人动画及人机共创。然而,该任务因两大挑战而研究不足:(1) 缺乏大规模精细3D指挥数据集;(2) 缺乏能同时支持长序列高质量高效生成的方法。为此,我们构建了以质量为导向的3D指挥动作采集流程,建立了包含约10小时数据的细粒度SMPL-X数据集CM-Data,据我们所知,这是首个且最大的公开3D指挥动作生成数据集。针对方法瓶颈,提出BiTDiff框架,基于BiMamba-Transformer混合模型架构与基于扩散的生成策略,结合人体运动学分解实现高质量动作合成。具体地,引入辅助物理一致性损失与手/身分离的前向运动学设计,提升精细动作建模能力,同时利用BiMamba实现内存高效的长序列建模,通过Transformer完成跨模态语义对齐。此外,支持无需训练的关节级动作编辑,推动下游人机交互设计。大量定量与定性实验表明,BiTDiff在CM-Data数据集上取得当前最优性能。代码将在录用后公开。

原文摘要 · Abstract (English)

3D conducting motion generation aims to synthesize fine-grained conductor motions from music, with broad potential in music education, virtual performance, digital human animation, and human-AI co-creation. However, this task remains underexplored due to two major challenges: (1) the lack of large-scale fine-grained 3D conducting datasets and (2) the absence of effective methods that can jointly support long-sequence generation with high quality and efficiency. To address the data limitation, we develop a quality-oriented 3D conducting motion collection pipeline and construct CM-Data, a fine-grained SMPL-X dataset with about 10 hours of conducting motion data. To the best of our knowledge, CM-Data is the first and largest public dataset for 3D conducting motion generation. To address the methodological limitation, we propose BiTDiff, a novel framework for 3D conducting motion generation, built upon a BiMamba-Transformer hybrid model architecture for efficient long-sequence modeling and a Diffusion-based generative strategy with human-kinematic decomposition for high-quality motion synthesis. Specifically, BiTDiff introduces auxiliary physical-consistency losses and a hand-/body-specific forward-kinematics design for better fine-grained motion modeling, while leveraging BiMamba for memory-efficient long-sequence temporal modeling and Transformer for cross-modal semantic alignment. In addition, BiTDiff supports training-free joint-level motion editing, enabling downstream human-AI interaction design. Extensive quantitative and qualitative experiments demonstrate that BiTDiff achieves state-of-the-art (SOTA) performance for 3D conducting motion generation on the CM-Data dataset. Code will be available upon acceptance.

3D生成动作合成扩散模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。