arXiv:2508.07650cs.RO2025-08AAAI被引 21

让机器人更懂模糊指令,3D空间感知提升操作成功率

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions

  • 用图结构建模3D物体与机械臂的拓扑关系,实现空间感知
  • 引入思维链推理模块,处理模糊指令并优化任务规划
  • 适合需要复杂环境交互的机器人任务,如家庭或工业场景

视觉-语言-动作模型在机器人操作中扮演关键角色,但现有方法在处理模糊语言指令和未知环境状态时存在明显局限,且感知能力受限于静态二维观察,缺乏对机器人与环境三维交互的建模能力。为此,本文提出GraphCoT-VLA,一种高效端到端模型。为增强对模糊指令的理解和任务规划能力,设计了融合高层任务理解、失败反馈与低层未来状态想象的结构化思维链推理模块。同时构建实时可更新的3D位姿-物体图,捕捉机器人关节与物体在三维空间中的配置关系及拓扑关联,提升交互理解能力。进一步引入丢弃式混合推理策略,实现高效控制输出。多组真实世界任务实验表明,GraphCoT-VLA在任务成功率与响应速度上显著优于现有方法,在开放环境和不确定指令下展现出强泛化性与鲁棒性。

原文摘要 · Abstract (English)

Vision-language-action models have emerged as a crucial paradigm in robotic manipulation. However, existing VLA models exhibit notable limitations in handling ambiguous language instructions and unknown environmental states. Furthermore, their perception is largely constrained to static two-dimensional observations, lacking the capability to model three-dimensional interactions between the robot and its environment. To address these challenges, this paper proposes GraphCoT-VLA, an efficient end-to-end model. To enhance the model's ability to interpret ambiguous instructions and improve task planning, we design a structured Chain-of-Thought reasoning module that integrates high-level task understanding and planning, failed task feedback, and low-level imaginative reasoning about future object positions and robot actions. Additionally, we construct a real-time updatable 3D Pose-Object graph, which captures the spatial configuration of robot joints and the topological relationships between objects in 3D space, enabling the model to better understand and manipulate their interactions. We further integrates a dropout hybrid reasoning strategy to achieve efficient control outputs. Experimental results across multiple real-world robotic tasks demonstrate that GraphCoT-VLA significantly outperforms existing methods in terms of task success rate and response speed, exhibiting strong generalization and robustness in open environments and under uncertain instructions.

机器人操作3D感知多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。