用分布强化学习处理人类干预信号不一致问题,提升机器人操控的鲁棒性。
GAINS: Leveraging Inconsistent Human Intervention Signals in Reinforcement Learning

- 引入分位数Q网络建模人类干预与稀疏奖励带来的回报不确定性。
- 在4个仿真和2个真实任务中,成功率比RLIF高22%,失败恢复提升43%。
- 适合需要人类实时修正的机器人操控场景,尤其在高控制频率下表现优。
通过人类干预纠正机器人操作策略对现实部署具有巨大潜力,但人类操作者在动作提供和干预时机上均存在固有缺陷。尽管动作不完美已被广泛研究,干预时机的不一致性在高控制频率下尤为突出,且仍缺乏深入探讨。本文提出GAINS框架,旨在利用不一致的人类干预信号进行强化学习。核心在于采用分位数Q网络的分布强化学习,建模由稀疏任务奖励和不一致干预引起的回报变异性。基于此分布表示,设计了一种保守探索策略,在人类修正下实现安全且样本高效的学习。我们在四个多样化的仿真操控任务及两个挑战性的真实世界场景中评估GAINS,相较于最先进的干预式方法,任务成功率提高22%,失败情况下的恢复成功率最高提升43%。结果凸显了建模人类不完美引发的回报变异性对干预式学习实际部署的重要性。
原文摘要 · Abstract (English)
Correcting robot manipulation policies through human intervention holds great promise for real-world deployment, yet human operators are inherently imperfect in both the actions they provide and the timing of their intervention signals. While the former has been extensively discussed in reinforcement learning (RL), the latter remains underexplored. At high control frequencies, human intervention signals are often delayed and inconsistent across time and state space. In this work, we present GAINS, a framework for leveraging inconsistent human intervention signals in RL. At the core of GAINS, we employ distributional RL with quantile Q-networks to model the return variability induced by sparse task rewards and inconsistent human interventions. Building on this distributional representation, we introduce a pessimistic exploration strategy that promotes safe and sample-efficient learning under human corrections. We evaluate GAINS on four diverse simulated manipulation tasks and two challenging real-world scenarios against state-of-the-art intervention-based methods. GAINS achieves a 22% higher task success rate than RLIF and improves recovery success by up to 43% in failure scenarios. These results highlight the importance of modeling return variability induced by human imperfection for real-world deployment of intervention-based learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。