arXiv:2510.07975cs.ROcs.AI2025-10被引 1

用数学化操作蓝图让机器人精准执行语言指令

Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation

  • 将视觉语言模型的语义理解转化为可执行的数学操作蓝图
  • 零样本泛化到多种复杂物体,模拟与真实环境均表现优异
  • 适合需要语义理解与物理控制衔接的研究者

让机器人在非结构化环境中实现精确且泛化的操作仍是具身智能的核心挑战。尽管视觉语言模型(VLM)在语义推理和任务规划方面表现出色,但其高层理解与实际操作所需的精确物理执行之间仍存在显著鸿沟。为此,我们提出GRACE框架,通过可执行分析概念(EAC)——一种数学定义的操作蓝图——将VLM推理具体化,编码物体功能、几何约束与操作语义。该方法构建结构化策略管线,将自然语言指令与视觉信息转化为实例化EAC,进而推导出抓取位姿、受力方向与物理可行运动轨迹。GRACE为高层指令理解与底层机器人控制提供统一可解释接口,有效实现语义-物理对齐下的精确泛化操作。大量实验表明,GRACE在模拟与真实世界中对多种铰接物体均实现强零样本泛化,无需任务特定训练。

原文摘要 · Abstract (English)

Enabling robots to perform precise and generalized manipulation in unstructured environments remains a fundamental challenge in embodied AI. While Vision-Language Models (VLMs) have demonstrated remarkable capabilities in semantic reasoning and task planning, a significant gap persists between their high-level understanding and the precise physical execution required for real-world manipulation. To bridge this "semantic-to-physical" gap, we introduce GRACE, a novel framework that grounds VLM-based reasoning through executable analytic concepts (EAC)-mathematically defined blueprints that encode object affordances, geometric constraints, and semantics of manipulation. Our approach integrates a structured policy scaffolding pipeline that turn natural language instructions and visual information into an instantiated EAC, from which we derive grasp poses, force directions and plan physically feasible motion trajectory for robot execution. GRACE thus provides a unified and interpretable interface between high-level instruction understanding and low-level robot control, effectively enabling precise and generalizable manipulation through semantic-physical grounding. Extensive experiments demonstrate that GRACE achieves strong zero-shot generalization across a variety of articulated objects in both simulated and real-world environments, without requiring task-specific training.

机器人操作视觉语言模型语义对齐零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。