arXiv:2606.23848cs.RO2026-06

用对抗方式让机械手弹琴更像人,不依赖专家数据

Enforcing Human-like Kinematics in Dexterous Piano Playing via Adversarial Posture Regularization

论文配图:Enforcing Human-like Kinematics in Dexterous Piano Playing via Adversarial Posture Regularization
图 1 · 摘自论文原文
  • 通过对抗机制匹配人类手部姿态分布,引导机械手自然动作
  • 在三个类人度指标上优于现有方法,视觉效果更真实
  • 仅需少量消费级设备采集的普通人演奏数据,无需专业标注

强化学习可训练双臂灵巧手在物理仿真中高精度弹奏钢琴,但对高自由度灵巧手而言,仅依赖任务奖励或逆运动学求解常导致动作不自然、关节过度伸展。本文提出对抗姿态正则化(APR),无需昂贵的歌曲对齐专家示范数据,而是利用少量普通人类演奏的非结构化数据。通过对抗目标使策略姿态分布逼近人类先验,从而鼓励更类人的手部形态。同时,我们使用消费级Meta Quest 3采集并发布了未结构化的钢琴演奏手势数据,并将关键运动信息迁移到Shadow Hand模型。实验表明,APR在所有三项类人度指标(cPSI、BSE、FAC)及视觉质量上均显著优于先前方法。

原文摘要 · Abstract (English)

Reinforcement learning can train bimanual dexterous hands to play piano in physics simulation with high note accuracy, but for high-DoF dexterous hands, relying solely on task rewards or IK inversion often leads to unnatural postures and joint overextension. We propose \textit{Adversarial Posture Regularization (APR)}. It avoids expensive, song-aligned expert demonstration data and instead uses a small amount of casual human playing data. By matching the distribution of the posture of the policy with the human prior through an adversarial objective, APR encourages more human-like hand shapes. Meanwhile, we collect and release unstructured hand motion data of piano playing using a consumer-grade Meta Quest 3, and retarget the key motion information to the Shadow Hand. Finally, we achieve significantly better performance than prior methods on all three human-likeness metrics (cPSI, BSE, and FAC) as well as in visual quality. Project repository: https://github.com/APRProject/APRPianist.

灵巧操作姿态正则化人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。