arXiv:2501.03841cs.RO2025-01CVPR被引 89

用物体功能空间构建操作指令,让机器人零样本完成复杂抓取。

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints

  • 以物体功能空间定义交互原语,连接语义理解与精确操作。
  • 无需微调大模型,在多种任务上实现零样本泛化表现。
  • 适合需要高泛化能力的仿真数据自动生成场景。

构建能在非结构化环境中通用操作的机器人系统仍是重大挑战。尽管视觉语言模型(VLM)在高层次常识推理上表现优异,但缺乏精细的3D空间理解能力以支持精准操作。对机器人数据微调VLM以生成视觉-语言-动作模型(VLA)虽有潜力,却受限于高昂的数据采集成本和泛化问题。为此,我们提出一种新型物体中心表征:利用物体的规范空间(由其功能可及性定义)作为结构化、语义明确的交互原语(如点与方向)描述方式,将VLM的常识推理转化为可执行的3D空间约束。在此基础上,我们设计了一种双闭环、开词汇的机器人操作框架:一个用于高层规划(原语重采样、交互渲染与VLM验证),另一个用于低层执行(6D姿态追踪)。该设计实现无需微调大模型的鲁棒实时控制。大量实验表明,该方法在多样化机器人操作任务中表现出强大的零样本泛化能力,展现出自动化大规模仿真数据生成的巨大潜力。

原文摘要 · Abstract (English)

The development of general robotic systems capable of manipulating in unstructured environments is a significant challenge. While Vision-Language Models(VLM) excel in high-level commonsense reasoning, they lack the fine-grained 3D spatial understanding required for precise manipulation tasks. Fine-tuning VLM on robotic datasets to create Vision-Language-Action Models(VLA) is a potential solution, but it is hindered by high data collection costs and generalization issues. To address these challenges, we propose a novel object-centric representation that bridges the gap between VLM's high-level reasoning and the low-level precision required for manipulation. Our key insight is that an object's canonical space, defined by its functional affordances, provides a structured and semantically meaningful way to describe interaction primitives, such as points and directions. These primitives act as a bridge, translating VLM's commonsense reasoning into actionable 3D spatial constraints. In this context, we introduce a dual closed-loop, open-vocabulary robotic manipulation system: one loop for high-level planning through primitive resampling, interaction rendering and VLM checking, and another for low-level execution via 6D pose tracking. This design ensures robust, real-time control without requiring VLM fine-tuning. Extensive experiments demonstrate strong zero-shot generalization across diverse robotic manipulation tasks, highlighting the potential of this approach for automating large-scale simulation data generation.

机器人操作视觉语言模型零样本泛化对象中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。