让机器人从示范中学习高层概念并泛化到新任务。
Bilevel Learning for Bilevel Planning
- 神经符号框架交替学习谓词的逻辑效果与函数表达。
- 在未见任务上平均成功率77%,远超现有方法(<35%)。
- 适合需要高阶抽象和泛化的机器人规划场景。
能从示范中学习的机器人不应仅模仿表面行为,而应理解演示中的高层概念,并将其泛化至新任务。双层规划是一种基于模型的分层方法,通过谓词(关系状态抽象)实现组合泛化。然而,以往方法依赖手工设计或形式简单的谓词,难以扩展到复杂、高维状态空间。为此,我们提出IVNTR,首个能直接从示范中学习神经谓词的双层规划方法。其核心创新是神经符号双层学习框架,模拟双层规划结构:符号学习谓词的“效果”与神经网络学习谓词的“函数”交替进行,相互提供指导。我们在六个不同的机器人规划领域评估IVNTR,证明其在抽象连续和高维状态方面的有效性。大多数现有方法泛化能力差(成功率<35%),而IVNTR在未见任务上平均达到77%成功率。此外,我们在移动操作臂上展示IVNTR,使其学会执行真实世界的移动操作任务,并泛化至包含新物体、新状态和更长任务时序的未见测试场景。研究结果凸显了基于抽象的学习与规划在实现高层泛化上的潜力。
原文摘要 · Abstract (English)
A robot that learns from demonstrations should not just imitate what it sees -- it should understand the high-level concepts that are being demonstrated and generalize them to new tasks. Bilevel planning is a hierarchical model-based approach where predicates (relational state abstractions) can be leveraged to achieve compositional generalization. However, previous bilevel planning approaches depend on predicates that are either hand-engineered or restricted to very simple forms, limiting their scalability to sophisticated, high-dimensional state spaces. To address this limitation, we present IVNTR, the first bilevel planning approach capable of learning neural predicates directly from demonstrations. Our key innovation is a neuro-symbolic bilevel learning framework that mirrors the structure of bilevel planning. In IVNTR, symbolic learning of the predicate "effects" and neural learning of the predicate "functions" alternate, with each providing guidance for the other. We evaluate IVNTR in six diverse robot planning domains, demonstrating its effectiveness in abstracting various continuous and high-dimensional states. While most existing approaches struggle to generalize (with <35% success rate), our IVNTR achieves an average of 77% success rate on unseen tasks. Additionally, we showcase IVNTR on a mobile manipulator, where it learns to perform real-world mobile manipulation tasks and generalizes to unseen test scenarios that feature new objects, new states, and longer task horizons. Our findings underscore the promise of learning and planning with abstractions as a path towards high-level generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。