用双流扩散模型同步生成钢琴双手动作,兼顾独立性与协调性。
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
- 双流架构分别建模左右手运动,通过独立噪声初始化实现解耦生成。
- 在多个指标上优于现有方法,显著提升动作协调性和手部特异性。
- 适合音乐生成、人机交互和虚拟演奏场景的研究者使用。
自动化合成协调的双手钢琴演奏面临巨大挑战,尤其在捕捉双手间复杂动作关系的同时保持各自独特的运动特征。本文提出一种双流神经框架,从音频输入生成同步的钢琴双手动作,解决了手部独立性与协调性建模的关键难题。框架引入两项创新:(i) 解耦的基于扩散的生成机制,通过双噪声初始化分别建模每只手的运动,同时共享位置条件;(ii) 手部协调非对称注意力(HCAA)机制,在去噪过程中抑制对称噪声,突出异步的手部特征,并自适应增强双手间的协调性。综合评估表明,该框架在多项指标上超越现有最先进方法。项目代码与演示见 https://monkek123king.github.io/S2C_page/。
原文摘要 · Abstract (English)
Automating the synthesis of coordinated bimanual piano performances poses significant challenges, particularly in capturing the intricate choreography between the hands while preserving their distinct kinematic signatures. In this paper, we propose a dual-stream neural framework designed to generate synchronized hand gestures for piano playing from audio input, addressing the critical challenge of modeling both hand independence and coordination. Our framework introduces two key innovations: (i) a decoupled diffusion-based generation framework that independently models each hand's motion via dual-noise initialization, sampling distinct latent noise for each while leveraging a shared positional condition, and (ii) a Hand-Coordinated Asymmetric Attention (HCAA) mechanism suppresses symmetric (common-mode) noise to highlight asymmetric hand-specific features, while adaptively enhancing inter-hand coordination during denoising. Comprehensive evaluations demonstrate that our framework outperforms existing state-of-the-art methods across multiple metrics. Our project is available at https://monkek123king.github.io/S2C_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。