arXiv:2509.21928cs.ROcs.AI2025-09

用场景图指导长程操作,让机器人更懂复杂任务的语义和步骤。

SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks

  • 用场景图结构化环境状态,连接高层规划与底层控制。
  • 在多个长程任务上达到当前最佳表现,成功完成未见任务。
  • 适合研究机器人规划、视觉导航与具身智能的学者。

解决长程操作任务仍面临根本挑战,这类任务涉及长时间的动作序列和复杂的物体交互,导致高层符号规划与底层连续控制之间存在显著差距。为弥合这一鸿沟,需具备鲁棒的长程任务规划和有效的目标条件操作能力。现有任务规划方法(包括传统和基于大语言模型的方法)普遍存在泛化能力差或语义推理稀疏的问题;而图像条件控制方法难以适应未见任务。为此,我们提出 SAGE 框架——一种面向长程操作任务的场景图感知引导与执行方法。SAGE 使用语义场景图作为场景状态的结构化表示,实现任务级语义推理与像素级视觉-运动控制的衔接,并支持可控生成准确的新子目标图像。该框架包含两个核心组件:(1) 基于场景图的任务规划器,利用视觉-语言模型(VLMs)和大语言模型(LLMs)解析环境并推理物理合理的状态转移序列;(2) 解耦的结构化图像编辑管道,通过图像修复与组合将每个目标子目标图转换为对应图像。大量实验表明,SAGE 在多个不同长程任务中达到当前最优性能。

原文摘要 · Abstract (English)

Successfully solving long-horizon manipulation tasks remains a fundamental challenge. These tasks involve extended action sequences and complex object interactions, presenting a critical gap between high-level symbolic planning and low-level continuous control. To bridge this gap, two essential capabilities are required: robust long-horizon task planning and effective goal-conditioned manipulation. Existing task planning methods, including traditional and LLM-based approaches, often exhibit limited generalization or sparse semantic reasoning. Meanwhile, image-conditioned control methods struggle to adapt to unseen tasks. To tackle these problems, we propose SAGE, a novel framework for Scene Graph-Aware Guidance and Execution in Long-Horizon Manipulation Tasks. SAGE utilizes semantic scene graphs as a structural representation for scene states. A structural scene graph enables bridging task-level semantic reasoning and pixel-level visuo-motor control. This also facilitates the controllable synthesis of accurate, novel sub-goal images. SAGE consists of two key components: (1) a scene graph-based task planner that uses VLMs and LLMs to parse the environment and reason about physically-grounded scene state transition sequences, and (2) a decoupled structural image editing pipeline that controllably converts each target sub-goal graph into a corresponding image through image inpainting and composition. Extensive experiments have demonstrated that SAGE achieves state-of-the-art performance on distinct long-horizon tasks.

机器人操作场景图长程规划视觉-语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。