arXiv:2411.18201cs.LGcs.AI2024-11中稿 · KDD被引 6

用符号推理提升长时序规划的数据效率与泛化能力。

Learning for Long-Horizon Planning via Neuro-Symbolic Abductive Imitation

  • 通过归纳推理将视觉输入转为符号表示,免去人工标注。
  • 构建多目标策略集成,基于符号逻辑动态选择最优动作。
  • 在多种长时序任务中显著提升数据效率和泛化性。

近期学习模仿方法在观察-动作空间内模仿规划方面表现良好,但在开放环境中的长时序任务仍受限。传统符号规划虽擅长长时序任务,依赖人类定义的符号空间,难以处理高维视觉输入等非符号观测。本文受归纳学习启发,提出新型框架ABIL(Abductive Imitation Learning),融合数据驱动学习与符号推理优势,实现长时序规划。具体地,利用归纳推理理解示范中的符号语义,并设计序列一致性原则解决感知与推理间的冲突;生成谓词候选以实现从原始观测到符号空间的自动映射,无需繁琐的谓词标注;在此基础上构建策略集成,其基础策略基于不同逻辑目标,由符号推理统一管理。实验表明,该方法能有效提取任务相关符号信息,辅助模仿学习;尤其在多个长时序任务中展现出显著更高的数据效率与泛化能力,验证其作为长时序规划解决方案的潜力。

原文摘要 · Abstract (English)

Recent learning-to-imitation methods have shown promising results in planning via imitating within the observation-action space. However, their ability in open environments remains constrained, particularly in long-horizon tasks. In contrast, traditional symbolic planning excels in long-horizon tasks through logical reasoning over human-defined symbolic spaces but struggles to handle observations beyond symbolic states, such as high-dimensional visual inputs encountered in real-world scenarios. In this work, we draw inspiration from abductive learning and introduce a novel framework \textbf{AB}ductive \textbf{I}mitation \textbf{L}earning (ABIL) that integrates the benefits of data-driven learning and symbolic-based reasoning, enabling long-horizon planning. Specifically, we employ abductive reasoning to understand the demonstrations in symbolic space and design the principles of sequential consistency to resolve the conflicts between perception and reasoning. ABIL generates predicate candidates to facilitate the perception from raw observations to symbolic space without laborious predicate annotations, providing a groundwork for symbolic planning. With the symbolic understanding, we further develop a policy ensemble whose base policies are built with different logical objectives and managed through symbolic reasoning. Experiments show that our proposal successfully understands the observations with the task-relevant symbolics to assist the imitation learning. Importantly, ABIL demonstrates significantly improved data efficiency and generalization across various long-horizon tasks, highlighting it as a promising solution for long-horizon planning. Project website: \url{https://www.lamda.nju.edu.cn/shaojj/KDD25_ABIL/}.

长时序规划符号推理模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。