让机器人在人机协作中自适应优化奖励与动作,提升操作鲁棒性。
EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

- 通过人类反馈动态更新奖励模型,适应场景变化。
- 用流匹配生成连贯动作块,减少动作不连续问题。
- 无需额外交互即可迁移视觉策略,适合真实场景部署。
人机协同强化学习(HIL-RL)使机器人从有限的真实交互中学习高接触力操作,但部署时面临三大耦合挑战:静态视觉奖励模型在场景变化下失效;独立采样动作导致时间上不一致的运动;基于视觉的策略对外观变化仍敏感。本文提出EvoHIL,一个统一框架,在分阶段的人机协同学习过程中同步优化奖励模型、动作生成器和视觉域。首先,自演化奖励(SER)利用人类确认的正样本和临时弱负样本更新成功分类器;其次,动作流稳定(AFS)通过流匹配生成时间连贯的动作片段,以已执行动作前缀和示范行为为依据更新策略;第三,保留感知离线微调在重放历史交互数据的同时,将AFS的演员-评论家模型锚定于先前行为,实现视觉域适应而无需额外机器人交互。在Franka FR3和SO-101机械臂上六个操作任务下,控制光照变化条件下,EvoHIL相较人机协同与模仿学习基线,显著提升任务成功率、与人类确认标签的一致性、运动平滑度及完成时间。
原文摘要 · Abstract (English)
Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision-based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process. First, self-evolving reward (SER) adapts the success classifier from human-confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention-aware offline fine-tuning replays relit interaction data while anchoring the AFS actor-critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO-101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。