arXiv:2504.10002cs.ROcs.LG2025-04ICRA被引 3

用低秩矩阵高效适配机器人行为偏好,避免遗忘原任务能力

FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions

  • 通过低秩矩阵扩展奖励模型,实现小样本偏好学习
  • 在仿真和真实机器人任务中均有效避免灾难性奖励遗忘
  • 适合数据稀缺场景下的机器人行为个性化调整

基于偏好的强化学习(PbRL)是适应预训练机器人行为风格的有效方法:在保持完成原始任务能力的同时,使机器人遵循人类用户偏好。然而,在机器人领域收集偏好数据往往困难且耗时。本文研究了在低偏好数据条件下的预训练机器人适应问题。我们发现,此类条件下现有适配方法易发生灾难性奖励遗忘(CRF),即更新后的奖励模型过度拟合新偏好,导致智能体无法执行原任务。为此,我们提出在原始奖励模型中引入少量参数(低秩矩阵)以建模偏好适应。评估结果表明,该方法可在多个仿真基准任务和真实机器人任务中,高效且有效地调整机器人行为以符合人类偏好。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task. However, collecting preferences for the adaptation process in robotics is often challenging and time-consuming. In this work we explore the adaptation of pre-trained robots in the low-preference-data regime. We show that, in this regime, recent adaptation approaches suffer from catastrophic reward forgetting (CRF), where the updated reward model overfits to the new preferences, leading the agent to become unable to perform the original task. To mitigate CRF, we propose to enhance the original reward model with a small number of parameters (low-rank matrices) responsible for modeling the preference adaptation. Our evaluation shows that our method can efficiently and effectively adjust robotic behavior to human preferences across simulation benchmark tasks and multiple real-world robotic tasks.

强化学习机器人偏好学习低秩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。