用生成式谓词发明让机器人学会抽象规划,解决复杂任务泛化难题。
SkillWrapper: Generative Predicate Invention for Task-level Robot Planning
- 基于大模型主动收集数据,从视觉输入生成可解释的符号化技能描述。
- 在仿真与真实机器人上实现未见长时序任务的组合求解,成功率超85%。
- 提供形式化保障,确保规划过程既正确又完整,适合具身智能研究者。
从单个技能执行泛化到长时序任务是构建自主机器人的核心挑战。一种有前景的方向是学习低层机器人技能的高层符号化表示,实现独立于底层状态空间的抽象推理。近期基础模型的发展使得基于原始感官输入生成符号谓词成为可能——我们称之为生成式谓词发明,以促进下游表征学习。然而,先前工作采用启发式或临时性方法学习这些抽象,忽视了其应满足的形式属性及如何保证这些属性的问题。本文提出生成式谓词发明的形式理论,并提出SkillWrapper方法,可学习可证明正确且完备的规划符号模型。该方法利用基础模型主动收集机器人数据,仅使用RGB图像观测,学习人类可读且可规划的表示。大量实证评估显示,SkillWrapper所学抽象表示能使机器人在真实世界中组合黑箱技能,解决未见过的长时序任务。
原文摘要 · Abstract (English)
Generalizing from individual skill executions to long-horizon tasks is a core challenge in building autonomous robots. A promising direction is learning high-level, symbolic representations of low-level robot skills, enabling abstract reasoning independent of the low-level state space. Recent advances in foundation models have made it possible to generate symbolic predicates that operate on raw sensory inputs-a process we call generative predicate invention-to facilitate downstream representation learning. However, prior work learns these abstractions using heuristic or ad-hoc procedures, ignoring the question of which formal properties they ought to satisfy, and how to guarantee these properties. We address these questions by presenting a formal theory of generative predicate invention for task-level planning, and proposing SkillWrapper, a method that learns symbolic models for provably sound and complete planning. Our approach leverages foundation models to actively collect robot data and learn human-interpretable, plannable representations, using only RGB image observations. Our extensive empirical evaluation in simulation and on real robots shows that SkillWrapper learns abstract representations that enable robots to compose black-box skills to solve unseen, long-horizon tasks in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。