arXiv:2504.10011cs.RO2025-04

用视觉语言模型指导机器人一次生成复杂动作,解决遮挡问题

KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks

  • 通过关键词选择和关键点对生成,将多模态指令转为可序列化的运动基元
  • 在模拟与真实环境中实现遮挡场景下刀切、蛋糕裱花等任务的精准动作生成
  • 适合需要快速理解指令并执行复杂动作的机器人应用

动态运动基元(DMPs)能以模块化参数编码平滑机器人动作,但难以融合视觉与语言等多模态输入。为充分发挥其潜力,需增强其处理多模态信息的能力。同时,我们致力于拓展其在物体导向任务中的一次性复杂动作生成能力,因执行过程中常发生遮挡(如蛋糕裱花时刀具遮挡、揉面时手部遮挡)。现有方法依赖视觉-语言模型(VLMs)进行高层语义理解,但缺乏生成低层运动细节的能力,仅能作为高层指令与底层控制之间的桥梁。为此,本文提出KeyMPs框架,结合VLMs与DMPs序列化机制:利用VLMs的高层推理能力进行‘关键词标记基元选择’,通过其空间感知能力生成用于调整运动尺度的‘关键点对’,从而实现基于单次多模态输入的一次性动作生成。我们在两个遮挡密集的任务上验证该方法:模拟与真实环境中的物体切割,以及模拟环境下的蛋糕裱花。实验结果表明,相比其他集成VLM支持的DMP方法,KeyMPs表现出更优性能。

原文摘要 · Abstract (English)

Dynamic Movement Primitives (DMPs) provide a flexible framework wherein smooth robotic motions are encoded into modular parameters. However, they face challenges in integrating multimodal inputs commonly used in robotics like vision and language into their framework. To fully maximize DMPs' potential, enabling them to handle multimodal inputs is essential. In addition, we also aim to extend DMPs' capability to handle object-focused tasks requiring one-shot complex motion generation, as observation occlusion could easily happen mid-execution in such tasks (e.g., knife occlusion in cake icing, hand occlusion in dough kneading, etc.). A promising approach is to leverage Vision-Language Models (VLMs), which process multimodal data and can grasp high-level concepts. However, they typically lack enough knowledge and capabilities to directly infer low-level motion details and instead only serve as a bridge between high-level instructions and low-level control. To address this limitation, we propose Keyword Labeled Primitive Selection and Keypoint Pairs Generation Guided Movement Primitives (KeyMPs), a framework that combines VLMs with sequencing of DMPs. KeyMPs use VLMs' high-level reasoning capability to select a reference primitive through \emph{keyword labeled primitive selection} and VLMs' spatial awareness to generate spatial scaling parameters used for sequencing DMPs by generalizing the overall motion through \emph{keypoint pairs generation}, which together enable one-shot vision-language guided motion generation that aligns with the intent expressed in the multimodal input. We validate our approach through experiments on two occlusion-rich tasks: object cutting, conducted in both simulated and real-world environments, and cake icing, performed in simulation. These evaluations demonstrate superior performance over other DMP-based methods that integrate VLM support.

运动生成多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。