arXiv:2505.06079cs.ROcs.CV2025-05ICRA被引 7

用少量示范+三网络互教,提升偏好强化学习抗噪声能力

TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations

  • 三模型并行训练,互相传递低损失偏好对作为知识
  • 仅需1-3个示范,噪声达40%时仍保持90%任务成功率
  • 适合需要少样本示范且标注不靠谱的机器人操控场景

基于人类或视觉语言模型(VLM)标注者的偏好反馈常存在噪声,严重制约偏好强化学习的性能。为此,我们提出TREND框架,将少量专家示范与三教学策略结合,有效缓解噪声影响。该方法同时训练三个奖励模型,每个模型将其小损失偏好对视为有用知识,并传递给其他网络用于参数更新。显著的是,本方法仅需1至3个专家示范即可实现高性能表现。我们在多种机器人操作任务上进行评估,即使在高达40%的噪声水平下,成功率仍可达90%,充分验证了其在处理噪声偏好反馈方面的强大鲁棒性。

原文摘要 · Abstract (English)

Preference feedback collected by human or VLM annotators is often noisy, presenting a significant challenge for preference-based reinforcement learning that relies on accurate preference labels. To address this challenge, we propose TREND, a novel framework that integrates few-shot expert demonstrations with a tri-teaching strategy for effective noise mitigation. Our method trains three reward models simultaneously, where each model views its small-loss preference pairs as useful knowledge and teaches such useful pairs to its peer network for updating the parameters. Remarkably, our approach requires as few as one to three expert demonstrations to achieve high performance. We evaluate TREND on various robotic manipulation tasks, achieving up to 90% success rates even with noise levels as high as 40%, highlighting its effective robustness in handling noisy preference feedback. Project page: https://shuaiyihuang.github.io/publications/TREND.

强化学习偏好学习机器人控制噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。