arXiv:2601.07821cs.ROcs.AI2026-01被引 10

让机器人在真实环境中学习时自动防错并自我修复,减少失败率73%

Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

  • 用世界模型+离线训练的恢复策略,提前识别并规避可能失败的操作
  • 真实场景实验中失败率降低73.1%,性能平均提升11.3%
  • 适合需要高可靠性的真实机器人操控任务,如工业搬运、家庭服务

基于深度强化学习的后训练算法可提升机器人在特定目标下的泛化性、准确性和鲁棒性。然而,在真实环境探索中不可避免地会出现需人工干预的失败(IR Failures,如洒水或打碎玻璃),阻碍该范式的实际应用。为此,我们提出失败感知的离线到在线强化学习(FARL),旨在最小化真实世界强化学习中的失败。我们构建了FailureBench基准,涵盖常见需人工干预的失败场景,并提出一种结合基于世界模型的安全评价器与离线训练的恢复策略的算法,以防止在线探索中的失败。大量仿真和真实世界实验表明,FARL在后训练阶段显著降低了IR Failures,同时提升了性能与泛化能力。在真实世界强化学习后训练中,FARL将IR Failures降低73.1%,性能平均提升11.3%。视频与代码见https://failure-aware-rl.github.io。

原文摘要 · Abstract (English)

Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen inevitably, hindering the practical deployment of such a paradigm. To tackle this, we introduce Failure-Aware Offline-to-Online Reinforcement Learning (FARL), a new paradigm minimizing failures during real-world reinforcement learning. We create FailureBench, a benchmark that incorporates common failure scenarios requiring human intervention, and propose an algorithm that integrates a world-model-based safety critic and a recovery policy trained offline to prevent failures during online exploration. Extensive simulation and real-world experiments demonstrate the effectiveness of FARL in significantly reducing IR Failures while improving performance and generalization during online reinforcement learning post-training. FARL reduces IR Failures by 73.1% while elevating performance by 11.3% on average during real-world RL post-training. Videos and code are available at https://failure-aware-rl.github.io.

强化学习机器人安全控制离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。