arXiv:2603.09542cs.RO2026-03被引 1

用符号编码+强化学习,让机器人更聪明地理解指令并自主探索新动作。

NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

  • 引入符号编码器提取视觉语言中的结构化动作基元。
  • 在单次训练下超越现有方法,零样本泛化能力更强。
  • 适合研究数据高效机器人控制与自主探索的学者。

视觉-语言-动作(VLA)模型旨在将指令与视觉场景关联,并生成机器人操作的动作序列。尽管取得进展,当前VLA模型仍面临学习可复用动作基元、减少对大规模数据和复杂架构依赖、拓展演示外探索能力等挑战。为此,我们提出一种基于在线强化学习(RL)的神经符号视觉-语言-动作(NS-VLA)框架。该框架引入符号编码器,将视觉与语言特征嵌入并提取结构化基元;利用符号求解器实现数据高效的动作排序;通过在线强化学习优化生成过程,实现广泛探索。在机器人操作基准测试中,NS-VLA在单次训练和数据扰动设置下均优于以往方法,同时展现出优异的零样本泛化能力、高数据效率及扩展的探索空间。代码已开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable primitives, reducing reliance on large-scale data and complex architectures, and enabling exploration beyond demonstrations. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework via online reinforcement learning (RL). It introduces a symbolic encoder to embedding vision and language features and extract structured primitives, utilizes a symbolic solver for data-efficient action sequencing, and leverages online RL to optimize generation via expansive exploration. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability, high data efficiency and expanded exploration space. Our code is available.

机器人控制符号学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。