arXiv:2608.21355cs.RO2026-08

让机器人通过人演示学会感知物体重量、摩擦和软硬,提升抓取适应性。

ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations

论文配图:ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
图 1 · 摘自论文原文
  • 基于人类视觉与触觉示范,联合建模物体物理属性
  • 对已知物体质量识别准确率97.2%,刚度误差仅5.51% MAPE
  • 可直接部署于机器人,显著提升跨类别抓取成功率

近期基于视觉的动作模型在复杂操作中表现优异,但很少利用显式的物体物理属性来调整策略。本文提出ViTacPhys,一个结合视觉-触觉建模与数据采集的框架,能从人类操作示范中估计物体质量、摩擦系数类别及连续刚度。该模型基于60个刚性与柔性物体的数据训练,采用时序视觉-触觉建模、跨注意力多模态融合,并引入来自视觉-语言模型的语义先验。在已见物体上,质量分类准确率达97.2%,摩擦系数分类准确率达98.8%,刚度平均绝对百分比误差(MAPE)为5.51%;在已知类别中的未见物体上,质量准确率87.5%,摩擦系数准确率97.5%,刚度MAPE为9.08%。通过少量机器人遥操作数据、机器人风格视频增强和匹配动作的人类示范,将模型迁移至机器人端,并作为在线模块实现自适应抓取。所获物理属性条件化策略在分布内物体上达到95.0%抓取成功率,分布外物体达83.4%。对成功抓取的分布外物体,其力信号更贴近人类遥操作,优于ACT基线。

原文摘要 · Abstract (English)

Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classes, together with continuous stiffness, from human manipulation demonstrations. Trained on data from 60 rigid and deformable objects, ViTacPhys combines temporal visual-tactile modeling, cross-attention multimodal fusion, and a semantic prior derived from a vision-language model. On seen objects, it achieves 97.2% mass classification accuracy, 98.8% friction-coefficient classification accuracy, and a stiffness mean absolute percentage error (MAPE) of 5.51%. On held-out objects from known categories, it achieves 87.5% mass accuracy, 97.5% friction-coefficient accuracy, and a stiffness MAPE of 9.08%. We transfer ViTacPhys from the human domain to the robot domain using limited robot teleoperation data, robot-style video augmentation, and human demonstrations with matched actions, and deploy it as an online module for adaptive grasping. The resulting physical-property-conditioned policy achieves total grasping success rates of 95.0% on in-distribution objects and 83.4% on out-of-distribution objects. For out-of-distribution objects successfully grasped by both methods, its force profiles are more consistent with human teleoperation than those produced by ACT. These results demonstrate the feasibility of explicitly estimating and conditioning on object physical properties for real-world adaptive grasping.

机器人抓取多模态学习物理属性感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。