让触觉模型学会判断动态物体的旋转滑动方向,提升机器人操作精度。
GeoTLM: Geometry-aware Tactile-Language Models for Contact Motion Orientation Reasoning of Dynamic Objects

- 通过可微分几何表示保留触觉剪切场的对称性结构
- 在新物体上旋转方向识别准确率提升14.6%,真实传感器下滑动方向识别提升16.2%
- 适用于需要精准接触感知的动态抓取与交互任务
当前触觉-语言模型(TLMs)在材料与纹理识别中表现良好,但在接触密集场景下难以理解动态物体的物理属性,如旋转与滑动方向。初步实验显示,Sparsh与AnyTouch2等主流模型在GelSight Mini触觉数据上的旋转方向推理能力较弱。为此,我们提出GeoTLM,一种基于几何先验的触觉-语言模型,用于动态接触事件感知。核心思想是在语言级推理前保留并结构化触觉剪切场的几何特性,而非将低分辨率触觉令牌强行嵌入脆弱的闭式物理算子。我们设计了轻量级(仅14k参数)的可微分几何表示(DGR),通过接触掩码引导的剪切场表示,并采用反称七区域池化,其设计基于旋转接触产生反称变形模式的物理直觉。在旋转方向与滑动方向推理两个任务上进行实验,结果表明,相较于无几何编码器的相同主干网络,GeoTLM在新物体上的旋转方向准确率提升14.6%,真实传感器下的滑动方向准确率提升16.2%。本工作为物理合理的触觉-语言推理开辟新路径,具有强动态物体理解与接触丰富机器人操作潜力。
原文摘要 · Abstract (English)
Modern tactile-language models (TLMs) have shown potential for robot learning tasks, such as material and texture recognition. However, for contact-rich scenarios, these TLMs struggle to understand the physical properties of dynamic objects, such as rotation and sliding directions. For instance, our preliminary experiments reveal that popular TLMs, such as Sparsh and AnyTouch2, exhibit weak performance on basic rotation direction reasoning from GelSight Mini tactile data. This surprising gap inspires us to explore a novel research question: Can we inject physically grounded geometric priors into TLMs to enable reliable contact orientation reasoning of dynamic object properties? To this end, we propose GeoTLM, a novel geometric representation-guided TLM for the perception of dynamic contact events. Our key idea is to preserve and structure tactile shear-field geometry before language-level reasoning, rather than forcing low-resolution tactile tokens into fragile closed-form physics operators. To achieve this, we propose a lightweight (only 14k parameters) yet novel Differentiable Geometric Representation (DGR). Specifically, DGR learns a contact-mask-guided representation in the shear field and aggregates it through an antisymmetric seven-region pooling design, motivated by the physical intuition that rotational contact produces antisymmetric deformation patterns. We conduct experiments on two representative tasks: rotation direction and sliding direction reasoning. Extensive experiments show that GeoTLM improves novel-object rotation accuracy by +14.6% and real-sensor sliding accuracy by +16.2% over the same backbone without the geometric encoder. Overall, our work paves a new way for physically grounded tactile-language reasoning, with strong potential for dynamic object understanding and contact-rich robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。