融合视觉声音触觉,让机械手学会轻柔抓取易碎物。
Learning Gentle Grasping Using Vision, Sound, and Touch
- 用声音信号判断抓取力度是否温柔,训练多模态模型预测抓握稳定性与柔和度。
- 在1500次实验中,多模态模型比纯视觉模型准确率高3.27%。
- 无需传感器校准或力建模,适合真实场景中抓取脆弱物体。
日常生活中,我们常需处理易损物品(如水果),抓取时应使用最小必要力而非最大力。本文提出结合视觉、触觉和听觉信号,学习稳定且轻柔的抓取与重抓动作。具体地,利用音频信号作为抓取柔和度的指示器,训练一个从原始多模态输入端到端预测未来抓握候选的稳定性和柔和度的动作条件模型,从而选择并执行最优动作。在1,500次多指机械手抓取试验中,该模型验证了其预测性能(比仅依赖视觉的模型高出3.27%准确率),并提供了行为可解释性。最终,真实世界实验表明,采用训练后的多模态模型,稳定且轻柔抓取成功率比纯视觉基线高出17%。本方法无需触觉传感器校准或解析力建模,显著降低工程实现难度。数据集与视频见https://lasr.org/research/gentle-grasping。
原文摘要 · Abstract (English)
In our daily life, we often encounter objects that are fragile and can be damaged by excessive grasping force, such as fruits. For these objects, it is paramount to grasp gently -- not using the maximum amount of force possible, but rather the minimum amount of force necessary. This paper proposes using visual, tactile, and auditory signals to learn to grasp and regrasp objects stably and gently. Specifically, we use audio signals as an indicator of gentleness during the grasping, and then train an end-to-end action-conditional model from raw visuo-tactile inputs that predicts both the stability and the gentleness of future grasping candidates, thus allowing the selection and execution of the most promising action. Experimental results on a multi-fingered hand over 1,500 grasping trials demonstrated that our model is useful for gentle grasping by validating the predictive performance (3.27% higher accuracy than the vision-only variant) and providing interpretations of their behavior. Finally, real-world experiments confirmed that the grasping performance with the trained multi-modal model outperformed other baselines (17% higher rate for stable and gentle grasps than vision-only). Our approach requires neither tactile sensor calibration nor analytical force modeling, drastically reducing the engineering effort to grasp fragile objects. Dataset and videos are available at https://lasr.org/research/gentle-grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。