arXiv:2412.18877cs.RO2024-12被引 1

用语言控制生成不重叠的机器人抓取目标位姿

Goal State Generation for Robotic Manipulation Based on Linguistically Guided Hybrid Gaussian Diffusion

  • 结合语言指令与混合高斯扩散模型生成目标位姿
  • 在10种马克杯、5种货架上测试,成功率最高且点云重叠减少90%以上
  • 适合需要精准位置控制的复杂抓取任务,如挂杯场景

在机器人操作任务中,为被操作物体生成指定目标状态对运动规划至关重要。例如挂杯子时,杯子必须位于挂钩附近的可行区域。以往方法虽可生成多个可行目标状态,但生成位置随机,缺乏控制,难以应对钩子已被占用或有特定操作目标等约束场景。此外,现实中杯子与支架频繁接触,端到端模型生成的目标状态常导致点云重叠,影响后续机械臂路径规划。为此,本文提出基于语言引导的混合高斯扩散(LHGD)网络生成目标状态,并引入基于重力覆盖系数的优化方法进行精修。为评估语言指定分布下的表现,我们收集了10种马克杯在5种不同货架共10个钩子上的多个可行目标状态,并准备5种未见杯子设计用于验证。实验表明,该方法在单模态、多模态及语言指定分布任务中均取得最高成功率,显著降低点云重叠,直接生成无碰撞目标状态,避免机械臂额外避障操作。

原文摘要 · Abstract (English)

In robotic manipulation tasks, achieving a designated target state for the manipulated object is often essential to facilitate motion planning for robotic arms. Specifically, in tasks such as hanging a mug, the mug must be positioned within a feasible region around the hook. Previous approaches have enabled the generation of multiple feasible target states for mugs; however, these target states are typically generated randomly, lacking control over the specific generation locations. This limitation makes such methods less effective in scenarios where constraints exist, such as hooks already occupied by other mugs or when specific operational objectives must be met. Moreover, due to the frequent physical interactions between the mug and the rack in real-world hanging scenarios, imprecisely generated target states from end-to-end models often result in overlapping point clouds. This overlap adversely impacts subsequent motion planning for the robotic arm. To address these challenges, we propose a Linguistically Guided Hybrid Gaussian Diffusion (LHGD) network for generating manipulation target states, combined with a gravity coverage coefficient-based method for target state refinement. To evaluate our approach under a language-specified distribution setting, we collected multiple feasible target states for 10 types of mugs across 5 different racks with 10 distinct hooks. Additionally, we prepared five unseen mug designs for validation purposes. Experimental results demonstrate that our method achieves the highest success rates across single-mode, multi-mode, and language-specified distribution manipulation tasks. Furthermore, it significantly reduces point cloud overlap, directly producing collision-free target states and eliminating the need for additional obstacle avoidance operations by the robotic arm.

机器人操作扩散模型目标生成语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。