首个钢琴演奏手部动作数据集,可生成逼真物理级演奏动画。
FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance

- 用多视角视频+迪斯克兰维亚钢琴传感器捕获15位顶级钢琴家的10小时3D手部动作。
- 结合扩散模型生成参考轨迹,再通过强化学习提升动作精度与泛化能力。
- 适合角色动画、虚拟现实和人机交互领域研究者使用。
钢琴演奏需要高度敏捷、精确且协调的手部控制,现有高保真手部运动模型在角色动画、具身人工智能、生物力学及虚拟/增强现实等领域具有广泛应用前景。本文构建了首个大规模数据集,包含15位顶尖钢琴家演奏153首古典乐曲的约10小时3D手部动作与音频数据。为实现自然表演捕捉,采用无标记多视角视频方案,借助先进姿态估计模型重建动作,并利用专用雅马哈迪斯克兰维亚钢琴获取的高分辨率MIDI按键数据通过逆运动学进行优化。基于该数据集,我们提出一个可合成非训练曲目物理合理手部动作的流程:结合模仿学习与强化学习,建立双侧手与琴键相互作用的物理驱动控制策略。针对大规模数据集带来的采样效率问题,采用扩散模型生成自然参考轨迹,提供高层级运动路径与指法信息;进一步利用音乐相似性从实录数据中检索相似动作以增强强化学习策略的精度。实验表明,所提方法能生成自然、灵巧且可泛化的演奏动作。
原文摘要 · Abstract (English)
Piano playing requires agile, precise, and coordinated hand control that stretches the limits of dexterity. Hand motion models with the sophistication to accurately recreate piano playing have a wide range of applications in character animation, embodied AI, biomechanics, and VR/AR. In this paper, we construct a first-of-its-kind large-scale dataset that contains approximately 10 hours of 3D hand motion and audio from 15 elite-level pianists playing 153 pieces of classical music. To capture natural performances, we designed a markerless setup in which motions are reconstructed from multi-view videos using state-of-the-art pose estimation models. The motion data is further refined via inverse kinematics using the high-resolution MIDI key-pressing data obtained from sensors in a specialized Yamaha Disklavier piano. Leveraging the collected dataset, we developed a pipeline that can synthesize physically-plausible hand motions for musical scores outside of the dataset. Our approach employs a combination of imitation learning and reinforcement learning to obtain policies for physics-based bimanual control involving the interaction between hands and piano keys. To solve the sampling efficiency problem with the large motion dataset, we use a diffusion model to generate natural reference motions, which provide high-level trajectory and fingering (finger order and placement) information. However, the generated reference motion alone does not provide sufficient accuracy for piano performance modeling. We then further augmented the data by using musical similarity to retrieve similar motions from the captured dataset to boost the precision of the RL policy. With the proposed method, our model generates natural, dexterous motions that generalize to music from outside the training dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。