arXiv:2509.08354cs.ROcs.AI2025-09中稿 · IEEE Transactions …被引 8

让机器人像人一样通过感知身体状态来抓取物体,尤其擅长处理变形物体。

Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration

  • 用数据手套捕捉人体关节级触觉与运动感知数据,实现人机统一数据格式。
  • 设计基于极坐标图结构的统一表征,提升不同手型间抓取策略的通用性。
  • 通过时空图网络预测关节状态,结合力-位混合映射生成机器人抓取指令。

触觉和本体感觉对人类灵巧操作至关重要,使我们能通过本体感觉-运动整合可靠抓取物体。尽管机器人手可获取此类反馈,但将感官信息直接映射为动作仍具挑战。本文提出一种手套辅助的触觉-运动感知-预测框架,基于模仿学习实现从人类自然操作到机器人执行的抓取技能迁移,并在泛化抓取任务中验证有效性,包括处理可变形物体。首先,采用数据手套采集关节级触觉与运动数据,该手套适用于人类与机器人手,支持跨场景自然人手示范的数据采集,确保原始数据格式一致,便于对人机抓取性能进行评估。其次,基于极坐标构建多模态输入的统一图结构表示,显式融合形态差异,增强不同示范者与机器人手间的兼容性。进一步提出触觉-运动时空图网络(TK-STGN),利用多维子图卷积与注意力机制的LSTM层,从图输入中提取时空特征,预测各关节的节点状态,并通过力-位混合映射转化为最终控制指令。

原文摘要 · Abstract (English)

Tactile and kinesthetic perceptions are crucial for human dexterous manipulation, enabling reliable grasping of objects via proprioceptive sensorimotor integration. For robotic hands, even though acquiring such tactile and kinesthetic feedback is feasible, establishing a direct mapping from this sensory feedback to motor actions remains challenging. In this paper, we propose a novel glove-mediated tactile-kinematic perception-prediction framework for grasp skill transfer from human intuitive and natural operation to robotic execution based on imitation learning, and its effectiveness is validated through generalized grasping tasks, including those involving deformable objects. Firstly, we integrate a data glove to capture tactile and kinesthetic data at the joint level. The glove is adaptable for both human and robotic hands, allowing data collection from natural human hand demonstrations across different scenarios. It ensures consistency in the raw data format, enabling evaluation of grasping for both human and robotic hands. Secondly, we establish a unified representation of multi-modal inputs based on graph structures with polar coordinates. We explicitly integrate the morphological differences into the designed representation, enhancing the compatibility across different demonstrators and robotic hands. Furthermore, we introduce the Tactile-Kinesthetic Spatio-Temporal Graph Networks (TK-STGN), which leverage multidimensional subgraph convolutions and attention-based LSTM layers to extract spatio-temporal features from graph inputs to predict node-based states for each hand joint. These predictions are then mapped to final commands through a force-position hybrid mapping.

机器人抓取模仿学习多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。