用对抗方式让机械手弹琴更像人,不依赖专家数据
Enforcing Human-like Kinematics in Dexterous Piano Playing via Adversarial Posture Regularization

- 通过对抗机制匹配人类手部姿态分布,引导机械手自然动作
- 在三个类人度指标上优于现有方法,视觉效果更真实
- 仅需少量消费级设备采集的普通人演奏数据,无需专业标注
强化学习可训练双臂灵巧手在物理仿真中高精度弹奏钢琴,但对高自由度灵巧手而言,仅依赖任务奖励或逆运动学求解常导致动作不自然、关节过度伸展。本文提出对抗姿态正则化(APR),无需昂贵的歌曲对齐专家示范数据,而是利用少量普通人类演奏的非结构化数据。通过对抗目标使策略姿态分布逼近人类先验,从而鼓励更类人的手部形态。同时,我们使用消费级Meta Quest 3采集并发布了未结构化的钢琴演奏手势数据,并将关键运动信息迁移到Shadow Hand模型。实验表明,APR在所有三项类人度指标(cPSI、BSE、FAC)及视觉质量上均显著优于先前方法。
原文摘要 · Abstract (English)
Reinforcement learning can train bimanual dexterous hands to play piano in physics simulation with high note accuracy, but for high-DoF dexterous hands, relying solely on task rewards or IK inversion often leads to unnatural postures and joint overextension. We propose \textit{Adversarial Posture Regularization (APR)}. It avoids expensive, song-aligned expert demonstration data and instead uses a small amount of casual human playing data. By matching the distribution of the posture of the policy with the human prior through an adversarial objective, APR encourages more human-like hand shapes. Meanwhile, we collect and release unstructured hand motion data of piano playing using a consumer-grade Meta Quest 3, and retarget the key motion information to the Shadow Hand. Finally, we achieve significantly better performance than prior methods on all three human-likeness metrics (cPSI, BSE, and FAC) as well as in visual quality. Project repository: https://github.com/APRProject/APRPianist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。