无需人类示范,机器人可自主演奏近千首钢琴曲。
Dexterous Robotic Piano Playing at Scale
- 用最优传输算法自动生成指法,无需演示即可学习演奏策略。
- 训练超2000个专用代理,构建百万级轨迹数据集RP1M++。
- 基于流匹配的Transformer模型,实现大规模钢琴演奏泛化能力。
赋予机器人手部类人灵巧性是机器人学长期目标。双手协同钢琴演奏尤为挑战:维度高、接触丰富且需快速精确控制。本文提出OmniPianist,首个通过可扩展、无演示学习完成近一千首乐曲演奏的智能体。方法包含三大核心组件:首先,基于最优传输(Optimal Transport, OT)设计自动指法策略,使智能体可从零开始自主发现高效演奏方案;其次,通过大规模强化学习训练超过2000个专属代理,每个代理专注于特定乐曲,聚合其经验形成名为RP1M++的数据集,包含逾一百万条机器人钢琴演奏轨迹;最后,采用流匹配变换器(Flow Matching Transformer)利用RP1M++进行大规模模仿学习,生成具备广泛曲目适应能力的OmniPianist智能体。大量实验与消融研究验证了方法的有效性与可扩展性,推动了灵巧机器人钢琴演奏的规模化发展。
原文摘要 · Abstract (English)
Endowing robot hands with human-level dexterity has been a long-standing goal in robotics. Bimanual robotic piano playing represents a particularly challenging task: it is high-dimensional, contact-rich, and requires fast, precise control. We present OmniPianist, the first agent capable of performing nearly one thousand music pieces via scalable, human-demonstration-free learning. Our approach is built on three core components. First, we introduce an automatic fingering strategy based on Optimal Transport (OT), allowing the agent to autonomously discover efficient piano-playing strategies from scratch without demonstrations. Second, we conduct large-scale Reinforcement Learning (RL) by training more than 2,000 agents, each specialized in distinct music pieces, and aggregate their experience into a dataset named RP1M++, consisting of over one million trajectories for robotic piano playing. Finally, we employ a Flow Matching Transformer to leverage RP1M++ through large-scale imitation learning, resulting in the OmniPianist agent capable of performing a wide range of musical pieces. Extensive experiments and ablation studies highlight the effectiveness and scalability of our approach, advancing dexterous robotic piano playing at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。