arXiv:2503.10546cs.ROcs.AI2025-03ICRA被引 6

用关键点统一动态学习与视觉提示,实现开放词汇机器人操作

KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation

  • 以关键点为桥梁,连接语言模型与动态规划
  • 支持自由语言指令下的多物体、可变形物体操作
  • 适合需要灵活交互的复杂机器人任务场景

随着大语言模型(LLMs)和视觉语言模型(VLMs)的发展,开放词汇机器人操作取得了显著进展。然而,许多现有方法忽视了物体动态特性,限制了其在复杂动态任务中的应用。本文提出KUDA,一个通过关键点统一动态学习与视觉提示的开放词汇操作框架,结合VLMs与基于学习的神经动力学模型。核心思想是:关键点目标描述既可被VLM理解,又能高效转化为模型预测的代价函数。给定语言指令与视觉观测后,KUDA首先在RGB图像上标注关键点,并调用VLM生成目标描述。这些抽象的关键点表示被转换为代价函数,再利用学习到的动力学模型优化,生成机器人轨迹。我们在多种任务上评估了KUDA,包括跨类别物体的自由语言指令、多物体交互以及可变形或颗粒状物体操作,验证了该框架的有效性。

原文摘要 · Abstract (English)

With the rapid advancement of large language models (LLMs) and vision-language models (VLMs), significant progress has been made in developing open-vocabulary robotic manipulation systems. However, many existing approaches overlook the importance of object dynamics, limiting their applicability to more complex, dynamic tasks. In this work, we introduce KUDA, an open-vocabulary manipulation system that integrates dynamics learning and visual prompting through keypoints, leveraging both VLMs and learning-based neural dynamics models. Our key insight is that a keypoint-based target specification is simultaneously interpretable by VLMs and can be efficiently translated into cost functions for model-based planning. Given language instructions and visual observations, KUDA first assigns keypoints to the RGB image and queries the VLM to generate target specifications. These abstract keypoint-based representations are then converted into cost functions, which are optimized using a learned dynamics model to produce robotic trajectories. We evaluate KUDA on a range of manipulation tasks, including free-form language instructions across diverse object categories, multi-object interactions, and deformable or granular objects, demonstrating the effectiveness of our framework. The project page is available at http://kuda-dynamics.github.io.

机器人操作视觉语言模型动态建模开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。