arXiv:2503.05064cs.ROcs.AI2025-03

用双层框架让机器人更精准地完成复杂装配任务

Perceiving, Reasoning, Adapting: A Dual-Layer Framework for VLM-Guided Precision Robotic Manipulation

  • 将任务分解为子动作,融合三维空间与二维拓扑信息
  • 实测在复杂装配中实现快速高精度操作,支持实时纠错
  • 适合需要精细动作规划的工业机器人场景

视觉语言模型(VLMs)在机器人操作中展现出巨大潜力,但在高速高精度的精细操作任务上仍面临挑战。现有方法虽擅长高层规划,却难以指导机器人完成精确的细小动作序列。为此,我们提出一种渐进式VLM规划算法,使机器人能够快速、精准且可纠错地完成精细操作。该方法将复杂任务分解为子动作,并维护三个关键数据结构:任务记忆结构、2D拓扑图和3D空间网络,实现高精度的空间-语义融合。这些组件在任务执行过程中持续积累并存储关键信息,为面向任务的VLM交互机制提供丰富上下文。由此,VLM能根据实时反馈动态调整引导策略,生成精确动作计划,并支持分步错误纠正。在复杂装配任务上的实验验证表明,该算法有效引导机器人在挑战性场景下快速精准完成精细操作,显著提升了机器人在精密任务中的智能水平。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) demonstrate remarkable potential in robotic manipulation, yet challenges persist in executing complex fine manipulation tasks with high speed and precision. While excelling at high-level planning, existing VLM methods struggle to guide robots through precise sequences of fine motor actions. To address this limitation, we introduce a progressive VLM planning algorithm that empowers robots to perform fast, precise, and error-correctable fine manipulation. Our method decomposes complex tasks into sub-actions and maintains three key data structures: task memory structure, 2D topology graphs, and 3D spatial networks, achieving high-precision spatial-semantic fusion. These three components collectively accumulate and store critical information throughout task execution, providing rich context for our task-oriented VLM interaction mechanism. This enables VLMs to dynamically adjust guidance based on real-time feedback, generating precise action plans and facilitating step-wise error correction. Experimental validation on complex assembly tasks demonstrates that our algorithm effectively guides robots to rapidly and precisely accomplish fine manipulation in challenging scenarios, significantly advancing robot intelligence for precision tasks.

机器人操作视觉语言模型精细控制动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。