让机器人自主纠错,减少人工干预57%还提升成功率8.6%
UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning

- 用未来动作价值预测判断探索是否无效
- 发现持续停滞就自动恢复,成功率达8.6%提升
- 适合需要高效人机协作的现实机器人任务
人在回路强化学习(HiL-RL)在真实世界机器人操作中表现出色,可通过人类指导在线优化策略。但现有框架依赖频繁人工干预以纠正无效探索,导致人力成本高、难以推广。为此,我们提出UniIntervene,一种智能体式干预模型,可自主检测无效探索并引导策略回归高价值状态,大幅减轻人工负担。具体而言,UniIntervene首先进行未来条件下的动作价值估计,预测当前动作的潜在后果及其价值,提供更稳定的进展信号;随后,时间价值风险判别器聚合近期价值动态,在价值持续停滞或下降时触发干预;此时,模型从历史干预经验中检索高价值恢复目标,并通过目标条件化的恢复策略生成可执行修正动作。该方法将干预从被动纠错转变为价值感知的主动恢复过程。在多种真实世界操纵任务上的实验证明,相较于最先进方法,UniIntervene平均成功率提升8.6%,人类干预次数减少57%。
原文摘要 · Abstract (English)
Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance. However, current HiL-RL frameworks remain intervention-intensive, relying on frequent human corrections to redirect the policy out of unproductive exploration, which incurs high labor cost and limits real-world scalability. To address this, we propose UniIntervene, an agentic intervention model that detects unproductive exploration and autonomously recovers the policy toward high-value states, taking over the bulk of interventions from human operators. Specifically, UniIntervene first performs future-conditioned action-value estimation, predicting the latent consequence of the current action and evaluating its induced value, which provides a more stable progress signal. Building on this, a temporal value-risk critic aggregates recent value dynamics and triggers intervention when the estimated value exhibits sustained stagnation or degradation. When intervention is required, UniIntervene retrieves a high-value recovery target from a memory of past intervention episodes and produces executable corrective actions through a goal-conditioned recovery policy. In this way, UniIntervene turns intervention from passive human correction into a value-aware recovery process for efficient real-world RL. Extensive experiments on diverse real-world manipulation tasks demonstrate that UniIntervene improves the average success rate by 8.6% while reducing human interventions by 57% relative to state-of-the-art HiL-RL baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。