arXiv:2511.04357cs.ROcs.CV2025-11被引 3

用图结构符号化表示演示,让机器人长任务规划更智能。

GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies

  • 用连续场景图提取人类示范的符号化动作表示
  • 实验证明可生成新规划域并支持连续执行多步动作
  • 适合需要长期规划的机器人自主学习场景

部署能从示范中学习新技能的自主机器人是现代机器人学的重要挑战。现有方法通常采用端到端模仿学习的视觉-语言-动作(VLA)模型或基于动作模型学习(AML)的符号化方法。前者缺乏高层符号规划能力,限制了其在长时序任务中的表现;后者则存在泛化与可扩展性不足的问题。本文提出一种新的神经符号框架GraSP-VLA,利用连续场景图表示从观察中生成人类示范的符号化动作表示。该表示在推理时用于生成新的规划域,并作为低层VLA策略的调度器,显著提升可连续执行的动作数量。实验表明,GraSP-VLA在自动规划域生成任务上具有有效性;真实世界实验进一步验证了其连续场景图表示在长时序任务中协调底层VLA策略的潜力。

原文摘要 · Abstract (English)

Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks.

机器人符号规划长时任务神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。