arXiv:2412.04835cs.ROcs.AI2024-12被引 11

用更少人类反馈让机器人视觉动作策略更符合用户偏好。

Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment

  • 仅通过人类偏好判断微调视觉编码器,再基于对齐特征生成奖励。
  • 实验表明只需5倍少的反馈数据即可完成策略对齐。
  • 适合希望降低人工标注成本的机器人研发人员。

视觉动作机器人策略在大规模数据集上预训练后,在多个机器人领域展现出巨大潜力。然而,如何将这些策略与终端用户偏好对齐仍具挑战,尤其当偏好难以明确表达时。尽管基于人类反馈强化学习(RLHF)在大语言模型等非具身领域取得成功,但在视觉动作策略对齐中因需大量人类反馈而受限。为此,我们提出表示对齐的偏好学习(RAPL),一种仅依赖观察的视觉奖励学习方法,显著减少所需的人类偏好反馈量。与传统RLHF不同,RAPL聚焦于通过人类反馈微调预训练视觉编码器,使其与用户视觉表征对齐,随后在该对齐空间中通过特征匹配构建密集视觉奖励。我们在X-Magical基准和Franka Panda机械臂操控仿真中验证了RAPL的有效性,结果表明其能高效学习符合人类偏好的奖励,并在不同机器人形态间实现泛化。最后,硬件实验对三个物体操作任务的预训练扩散策略进行对齐,发现RAPL仅需5倍少的真实人类偏好数据即可完成微调,为最小化人类反馈、最大化机器人策略对齐迈出关键一步。

原文摘要 · Abstract (English)

Visuomotor robot policies, increasingly pre-trained on large-scale datasets, promise significant advancements across robotics domains. However, aligning these policies with end-user preferences remains a challenge, particularly when the preferences are hard to specify. While reinforcement learning from human feedback (RLHF) has become the predominant mechanism for alignment in non-embodied domains like large language models, it has not seen the same success in aligning visuomotor policies due to the prohibitive amount of human feedback required to learn visual reward functions. To address this limitation, we propose Representation-Aligned Preference-based Learning (RAPL), an observation-only method for learning visual rewards from significantly less human preference feedback. Unlike traditional RLHF, RAPL focuses human feedback on fine-tuning pre-trained vision encoders to align with the end-user's visual representation and then constructs a dense visual reward via feature matching in this aligned representation space. We first validate RAPL through simulation experiments in the X-Magical benchmark and Franka Panda robotic manipulation, demonstrating that it can learn rewards aligned with human preferences, more efficiently uses preference data, and generalizes across robot embodiments. Finally, our hardware experiments align pre-trained Diffusion Policies for three object manipulation tasks. We find that RAPL can fine-tune these policies with 5x less real human preference data, taking the first step towards minimizing human feedback while maximizing visuomotor robot policy alignment.

机器人对齐视觉奖励少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。