机器人从失败中学习抽象概念,实现自主避错与规划。
Recover, Discover, Plan: Learning Skills and Concepts from Robot Failures

- 通过失败-恢复经验逐步发现并优化状态抽象关系
- 在4个模拟环境中解决未见过的长程任务,性能超基线50%以上
- 适合需要自适应避错的复杂物理世界机器人应用
智能机器人不仅要能从失败中恢复,还需获取避免未来失败所需的抽象知识。尽管强化学习可学习反应式恢复行为,但为每种故障模式单独训练策略效率极低。本文提出ReSYNC——首个从失败-恢复经验中渐进发现并精炼状态抽象(关系谓词)的方法,以支持抽象规划。不同于纯反应式方法,ReSYNC通过增量双学习过程联合学习技能与概念:在技能学习阶段,机器人使用强化学习学习从训练任务中遇到的故障中恢复;在概念学习阶段,机器人发现新的关系谓词并优化其抽象规划模型,以解释和泛化所学的恢复行为。这种交互使ReSYNC能将训练中观察到的局部恢复转化为测试时的全局故障规避。在四个模拟领域中,我们展示了ReSYNC持续扩展与精炼其抽象库的能力,使其能够解决长程、此前未见的问题,性能超过强基线50%以上。此外,我们验证了ReSYNC的模拟到现实迁移能力,其在真实世界中执行非抓取操纵技能,并通过抽象规划泛化至未见场景。总体而言,ReSYNC代表了迈向能在物理世界中自主获取抽象知识、实现可扩展故障感知规划的重要一步。
原文摘要 · Abstract (English)
Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoid them in the future. While reinforcement learning (RL) can learn reactive recovery behaviors, training a separate policy for every distinct failure mode is highly inefficient. We introduce Recovery-Driven Synthesis of Relational Concepts (ReSYNC), the first approach that progressively discovers and refines state abstractions (relational predicates) from failure-recovery experience to support abstract planning. Unlike purely reactive methods, ReSYNC jointly learns skills and concepts through an incremental dual-learning process. In the skill-learning phase, the robot uses RL to learn to recover from failures seen in training tasks. In the concept-learning phase, the robot discovers new relational predicates and refines its abstract planning model to explain and generalize the learned recovery behaviors. This interaction enables ReSYNC to convert local recoveries seen during training into global failure avoidance at test time. Across four simulated domains, we show that ReSYNC's ability to continually expand and refine its abstraction library allows it to solve long-horizon, previously unseen problems, outperforming strong baselines by over 50%. Additionally, we demonstrate sim-to-real transfer of ReSYNC, where it performs real-world non-prehensile manipulation skills and generalizes to unseen scenarios through abstract planning. Overall, ReSYNC represents a significant step toward robots that autonomously acquire abstractions for scalable, failure-aware planning in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。