用扩散模型生成精准且富有表现力的机器人钢琴演奏动作。
PANDORA: Diffusion Policy Learning for Dexterous Robotic Piano Playing
- 基于条件U-Net与FiLM全局条件,迭代去噪生成平滑高维动作序列。
- 融合任务精度、音频保真度与大语言模型反馈,实现动态奖励调整。
- 在ROBOPIANIST上超越基线,适合追求音乐表现力的机器人控制研究者。
我们提出PANDORA,一种专为灵巧机器人钢琴演奏设计的基于扩散的概率策略学习框架。该方法采用带FiLM全局条件的条件U-Net架构,通过迭代去噪将噪声动作序列转化为平滑的高维轨迹。为实现精确按键与富有表现力的音乐演奏,我们设计了复合奖励函数,整合任务特定精度、音频保真度以及来自大语言模型(LLM)的高层语义反馈。该LLM代理评估音乐表现力与风格细节,支持手部特异性奖励动态调节。进一步引入残差逆运动学精修策略,PANDORA在ROBOPIANIST环境中达到当前最优性能,显著优于基线,在精度与表现力上均有提升。消融实验验证了扩散去噪与LLM驱动语义反馈对增强机器人音乐性的关键作用。视频展示见:https://taco-group.github.io/PANDORA
原文摘要 · Abstract (English)
We present PANDORA, a novel diffusion-based policy learning framework designed specifically for dexterous robotic piano performance. Our approach employs a conditional U-Net architecture enhanced with FiLM-based global conditioning, which iteratively denoises noisy action sequences into smooth, high-dimensional trajectories. To achieve precise key execution coupled with expressive musical performance, we design a composite reward function that integrates task-specific accuracy, audio fidelity, and high-level semantic feedback from a large language model (LLM) oracle. The LLM oracle assesses musical expressiveness and stylistic nuances, enabling dynamic, hand-specific reward adjustments. Further augmented by a residual inverse-kinematics refinement policy, PANDORA achieves state-of-the-art performance in the ROBOPIANIST environment, significantly outperforming baselines in both precision and expressiveness. Ablation studies validate the critical contributions of diffusion-based denoising and LLM-driven semantic feedback in enhancing robotic musicianship. Videos available at: https://taco-group.github.io/PANDORA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。