arXiv:2504.15595cs.RO2025-04被引 5

用视觉触觉融合让机器人稳稳抓捏易变形物体,不丢不破。

Grasping Deformable Objects via Reinforcement Learning with Cross-Modal Attention to Visuo-Tactile Inputs

  • 通过跨模态注意力融合视觉与触觉信息,提升感知能力。
  • 在多种未见物体和动作下表现优于早期/晚期融合方法。
  • 适合研究机器人灵巧操作与多模态感知的学者参考。

本文研究使用机械夹爪抓取具有柔性外壳的可变形物体的问题。这类物体质心动态变化且易破裂,导致机器人难以生成合适的控制指令以避免掉落或损坏。多模态传感数据可提供全局信息(如形状、姿态)和局部接触信息(如压力),二者互补但融合困难。本文提出基于深度强化学习(DRL)的方法,从视觉与触觉输入中生成夹爪控制指令。模型在编码器中引入跨模态注意力模块,并通过强化学习代理的损失函数进行自监督训练。实验表明,该方法在不同环境(包括未见过的机器人动作与物体)下均优于早期和晚期融合方法,证明跨模态注意力有效提升了感知表示能力。

原文摘要 · Abstract (English)

We consider the problem of grasping deformable objects with soft shells using a robotic gripper. Such objects have a center-of-mass that changes dynamically and are fragile so prone to burst. Thus, it is difficult for robots to generate appropriate control inputs not to drop or break the object while performing manipulation tasks. Multi-modal sensing data could help understand the grasping state through global information (e.g., shapes, pose) from visual data and local information around the contact (e.g., pressure) from tactile data. Although they have complementary information that can be beneficial to use together, fusing them is difficult owing to their different properties. We propose a method based on deep reinforcement learning (DRL) that generates control inputs of a simple gripper from visuo-tactile sensing information. Our method employs a cross-modal attention module in the encoder network and trains it in a self-supervised manner using the loss function of the RL agent. With the multi-modal fusion, the proposed method can learn the representation for the DRL agent from the visuo-tactile sensory data. The experimental result shows that cross-modal attention is effective to outperform other early and late data fusion methods across different environments including unseen robot motions and objects.

机器人抓取多模态感知强化学习触觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。