用少量示范实现长程任务规划与执行,兼具高效与可解释性。
Few-Shot Neuro-Symbolic Imitation Learning for Long-Horizon Planning and Acting
- 结合符号逻辑与神经网络,从少数示范中学习高层任务结构
- 仅需5次示范即在6个场景中实现零样本与少样本泛化
- 适合需要可靠决策过程的机器人长程任务应用
模仿学习使智能系统能在极少监督下掌握复杂行为。然而,现有方法多聚焦短时技能,依赖大量数据,难以应对长程任务或任务变化与分布偏移。本文提出一种新型神经符号框架,联合学习连续控制策略与符号领域抽象,仅需少量技能示范。该方法将高层任务结构抽象为图结构,通过答案集编程求解器发现符号规则,并利用扩散策略模仿学习训练底层控制器。高层智能体筛选任务相关信 息,使每个控制器仅关注最小观测与动作空间。基于图的神经符号框架能捕捉复杂状态转移,包括非空间与时间关系,而这些是数据驱动方法在有限示范下常难以发现的。我们在六个领域验证了该方法:包含四种机械臂的堆叠、厨房、装配与汉诺塔环境,以及具两个场景的自动叉车域。结果表明,该方法具有高数据效率(最少5次示范)、强零样本与少样本泛化能力,且决策过程可解释。
原文摘要 · Abstract (English)
Imitation learning enables intelligent systems to acquire complex behaviors with minimal supervision. However, existing methods often focus on short-horizon skills, require large datasets, and struggle to solve long-horizon tasks or generalize across task variations and distribution shifts. We propose a novel neuro-symbolic framework that jointly learns continuous control policies and symbolic domain abstractions from a few skill demonstrations. Our method abstracts high-level task structures into a graph, discovers symbolic rules via an Answer Set Programming solver, and trains low-level controllers using diffusion policy imitation learning. A high-level oracle filters task-relevant information to focus each controller on a minimal observation and action space. Our graph-based neuro-symbolic framework enables capturing complex state transitions, including non-spatial and temporal relations, that data-driven learning or clustering techniques often fail to discover in limited demonstration datasets. We validate our approach in six domains that involve four robotic arms, Stacking, Kitchen, Assembly, and Towers of Hanoi environments, and a distinct Automated Forklift domain with two environments. The results demonstrate high data efficiency with as few as five skill demonstrations, strong zero- and few-shot generalizations, and interpretable decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。