arXiv:2609.06009cs.RO2026-09

让机器人从人类避开的动作中学习安全策略,减少干预和失败。

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

论文配图:How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation
图 1 · 摘自论文原文
  • 用人类干预信号训练预测模型,提前识别高风险状态。
  • 在16自由度机械手上实现96.7%抓取成功率,干预次数减少84%。
  • 适合需要安全、高效训练复杂机械臂的科研与工业场景。

多指灵巧操作仍是现实世界强化学习的前沿挑战,源于高维动作空间和硬件故障的高昂代价。尽管人机协同强化学习允许操作员在故障前介入,但现有方法常将干预视为被动纠正,忽略了其中蕴含的安全信号。本文提出WHIRL框架,将二元人类干预转化为前向预测信号,实现主动风险规避。该方法基于一种干预感知的潜在世界模型,包含四个预测头:动力学、奖励、终止以及一个新颖的逐状态干预概率头,可预测未来状态的人类接管概率。该头为策略提供风险调节项,引导其避开易引发干预的区域,建模操作员的安全阈值。我们在16自由度的LEAP手在多种任务上进行评估,涵盖凸形与非规则物体抓取、棱柱体操作及长时序多阶段任务。结果表明,预测性风险调节使系统在复杂抓取任务中达到96.7%的成功率,同时在步数加权下将操作员干预负担降低最多84%。本工作通过连接人类直觉与预测建模,为真实世界中复杂灵巧智能体的训练提供了实用的安全增强方案,有效减轻操作疲劳与硬件风险。

原文摘要 · Abstract (English)

Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the high-dimensional action space and the prohibitive cost of hardware failures. While human-in-the-loop (HIL) RL allows operators to intervene before failures occur, current pipelines often treat these interventions as reactive corrections, discarding the rich safety signal inherent in the operator's decision to take control. In this paper, we ask: How can we learn from what a human would avoid? We present WHIRL, a safety-aware RL framework that transforms binary human interventions into forward-predictive signals for proactive risk avoidance. Our approach centers on an intervention-aware latent world model with four prediction heads: dynamics, reward, termination, and a novel per-state intervention-probability head that learns to predict the likelihood of a human takeover at future states. This head provides an actor-side risk-shaping term that discourages the policy from entering "intervention-prone" regions, modeling the operator's internal safety threshold. We evaluate our framework on a 16-DoF LEAP Hand across tasks spanning convex and irregular object grasping, prismatic manipulation, and long-horizon multi-stage tasks. Our results show that predictive risk-shaping enables the system to achieve a 96.7 percent success rate on complex grasping tasks while reducing the operator intervention burden by up to 84 percent in step-weighted terms. By closing the loop between human intuition and predictive world modeling, this work provides a practical safety-aware recipe for training complex dexterous agents in the real world while reducing operator fatigue and hardware-risk exposure.

强化学习灵巧操作人机协作安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。