arXiv:2505.02166cs.RO2025-05被引 8

用简单视觉提示指导机器人完成复杂操作任务。

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

  • 为每个关键帧生成2D视觉提示,明确表达动作与目标。
  • 模型能准确预测SE(3)空间中的接触位姿与运动方向。
  • 适合需要灵活、可解释性操作的机器人场景。

在机器人任务中,目标可通过语言、图像或视频等多种模态传达。然而,自然语言可能存在歧义,而图像或视频又可能提供过于详细的规格。为此,我们提出CrayonRobo,一种基于多模态提示的物体中心型视觉-语言-动作模型。针对任务序列中的每个关键帧,该方法允许手动或自动在RGB图像上叠加简洁且富有表现力的2D视觉提示,以明确表达任务目标,如末端执行器位姿和接触后的运动方向。我们设计了一种训练策略,使模型能够理解这些视觉-语言提示,并在SE(3)空间中预测对应的接触位姿与运动方向。通过依次执行所有关键帧步骤,模型可完成长时程任务。该方法不仅帮助模型显式理解任务目标,还通过可解释的提示提升了对未见任务的鲁棒性。我们在仿真与真实环境中评估了该方法,验证了其强大的操纵能力。

原文摘要 · Abstract (English)

In robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To tackle these challenges, we introduce CrayonRobo that leverages comprehensive multi-modal prompts that explicitly convey both low-level actions and high-level planning in a simple manner. Specifically, for each key-frame in the task sequence, our method allows for manual or automatic generation of simple and expressive 2D visual prompts overlaid on RGB images. These prompts represent the required task goals, such as the end-effector pose and the desired movement direction after contact. We develop a training strategy that enables the model to interpret these visual-language prompts and predict the corresponding contact poses and movement directions in SE(3) space. Furthermore, by sequentially executing all key-frame steps, the model can complete long-horizon tasks. This approach not only helps the model explicitly understand the task objectives but also enhances its robustness on unseen tasks by providing easily interpretable prompts. We evaluate our method in both simulated and real-world environments, demonstrating its robust manipulation capabilities.

机器人操作多模态提示视觉引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。