arXiv:2503.02748cs.RO2025-03被引 1

用语义关键点连接视觉语言模型与运动基元,实现精准机器人操作。

Bridging VLM and KMP: Enabling Fine-grained robotic manipulation via Semantic Keypoints Representation

  • 通过语义关键点约束传递决策信息,实现跨框架低失真交互。
  • 在真实复杂场景中验证,支持复杂轨迹形状保持的精细操作。
  • 适合需要灵活适应与高精度控制的机器人任务研发人员。

从早期运动基元(MP)技术到现代视觉语言模型(VLM),自主操作一直是机器人领域的核心课题。作为两种极端,基于VLM的方法强调零样本与自适应操作,但在细粒度规划上表现不佳;而基于MP的方法擅长精确轨迹泛化,却缺乏决策能力。为此,我们提出VL-MP,通过低失真决策信息传输桥,将VLM与核化运动基元(KMP)结合,实现在模糊情境下的细粒度机器人操作。其关键在于利用语义关键点约束精确表示任务决策参数,从而生成更精准的任务参数。此外,我们引入局部轨迹特征增强的KMP以支持VL-MP,实现了复杂轨迹的形状保持。在复杂真实环境中的大量实验验证了VL-MP在自适应与细粒度操作方面的有效性。

原文摘要 · Abstract (English)

From early Movement Primitive (MP) techniques to modern Vision-Language Models (VLMs), autonomous manipulation has remained a pivotal topic in robotics. As two extremes, VLM-based methods emphasize zero-shot and adaptive manipulation but struggle with fine-grained planning. In contrast, MP-based approaches excel in precise trajectory generalization but lack decision-making ability. To leverage the strengths of the two frameworks, we propose VL-MP, which integrates VLM with Kernelized Movement Primitives (KMP) via a low-distortion decision information transfer bridge, enabling fine-grained robotic manipulation under ambiguous situations. One key of VL-MP is the accurate representation of task decision parameters through semantic keypoints constraints, leading to more precise task parameter generation. Additionally, we introduce a local trajectory feature-enhanced KMP to support VL-MP, thereby achieving shape preservation for complex trajectories. Extensive experiments conducted in complex real-world environments validate the effectiveness of VL-MP for adaptive and fine-grained manipulation.

机器人操作视觉语言模型运动基元语义关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。