arXiv:2603.25903cs.RO2026-03被引 1

让机器人从视觉动作数据中自动学出可解释的决策结构

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories

  • 用自适应聚类和L*算法从演示数据中提取离散状态机
  • 在低数据下比顶尖端到端模型性能高27%
  • 适合需要可解释性和长程规划的机器人任务

将机器人学习扩展到长时序任务仍是重大挑战。传统端到端策略缺乏长期推理所需的结构先验,而传统神经符号方法依赖手工设计的符号先验。为此,我们提出ENAP(Emergent Neural Automaton Policy),一种允许从视觉动作演示中自适应生成神经符号策略的框架。具体而言,我们首先利用自适应聚类和L*算法的扩展版本,从视觉动作数据中推断出一个梅利状态机,作为可解释的高层规划器,捕捉潜在的任务模式。随后,该离散结构引导一个低层反应式残差网络通过行为克隆学习精确的连续控制。通过显式建模离散转移与连续残差,ENAP在无需任务特定标签的情况下实现了高样本效率和可解释性。在复杂操作和长时序任务上的大量实验表明,ENAP在低数据环境下比最先进端到端视觉语言动作(VLA)策略性能最高提升27%,同时提供机器人意图的结构化表示(图1)。

原文摘要 · Abstract (English)

Scaling robot learning to long-horizon tasks remains a formidable challenge. While end-to-end policies often lack the structural priors needed for effective long-term reasoning, traditional neuro-symbolic methods rely heavily on hand-crafted symbolic priors. To address the issue, we introduce ENAP (Emergent Neural Automaton Policy), a framework that allows a bi-level neuro-symbolic policy adaptively emerge from visuomotor demonstrations. Specifically, we first employ adaptive clustering and an extension of the L* algorithm to infer a Mealy state machine from visuomotor data, which serves as an interpretable high-level planner capturing latent task modes. Then, this discrete structure guides a low-level reactive residual network to learn precise continuous control via behavior cloning (BC). By explicitly modeling the task structure with discrete transitions and continuous residuals, ENAP achieves high sample efficiency and interpretability without requiring task-specific labels. Extensive experiments on complex manipulation and long-horizon tasks demonstrate that ENAP outperforms state-of-the-art (SoTA) end-to-end VLA policies by up to 27% in low-data regimes, while offering a structured representation of robotic intent (Fig. 1).

机器人学习神经符号可解释性长程规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。