用人类偏好微调扩散策略,让机器人行为更贴合新需求。
FDPP: Fine-tune Diffusion Policy with Human Preference
- 通过偏好学习构建奖励函数,指导策略微调。
- 在保持原任务性能基础上,成功适配新偏好。
- 引入KL正则化防止过拟合,保留初始能力。
从人类示范中进行模仿学习使机器人能够执行复杂操作任务,近年来取得显著进展。然而,这类方法往往难以适应新偏好或环境变化。为此,我们提出基于人类偏好的扩散策略微调方法(FDPP)。FDPP通过偏好学习构建奖励函数,并利用强化学习对预训练策略进行微调,使策略在完成原任务的同时与新的人类偏好对齐。在多种机器人任务和偏好设置下的实验表明,FDPP能有效定制策略行为而不影响性能。此外,我们在微调过程中引入Kullback-Leibler(KL)正则化,防止过拟合,有助于保持初始策略的泛化能力。
原文摘要 · Abstract (English)
Imitation learning from human demonstrations enables robots to perform complex manipulation tasks and has recently witnessed huge success. However, these techniques often struggle to adapt behavior to new preferences or changes in the environment. To address these limitations, we propose Fine-tuning Diffusion Policy with Human Preference (FDPP). FDPP learns a reward function through preference-based learning. This reward is then used to fine-tune the pre-trained policy with reinforcement learning (RL), resulting in alignment of pre-trained policy with new human preferences while still solving the original task. Our experiments across various robotic tasks and preferences demonstrate that FDPP effectively customizes policy behavior without compromising performance. Additionally, we show that incorporating Kullback-Leibler (KL) regularization during fine-tuning prevents over-fitting and helps maintain the competencies of the initial policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。