构建动态触觉感知新基准,让机器人更懂物体接触时的细微变化。
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception
- 提出分层触觉数据集ToucHD,涵盖动作、交互与力反馈
- AnyTouch 2模型统一学习物体属性与力感知的动态变化
- 适用于多类光学触觉传感器,提升复杂操作能力
真实世界中丰富的抓取任务要求机器人感知时间序列触觉信号,捕捉微小表面形变,并推断物体属性及受力动态。尽管光学触觉传感器能提供此类丰富信息,现有数据集和模型仍受限:主要关注物体层面属性(如材质),而忽视物理交互中的细粒度动态特征。为此,我们提出一套系统化的动态感知能力层级,指导数据采集与模型设计。为填补动态数据空白,我们构建了大规模分层触觉数据集ToucHD,覆盖触觉原子动作、真实操作及触-力配对数据。ToucHD从数据层面支持多层次感知能力。在此基础上,提出AnyTouch 2——一种通用光学触觉表示学习框架,可适配多种传感器,统一实现物体级理解与力感知的精细动态建模。该框架在帧间捕捉像素级与动作特异性形变,并显式建模物理力动态,从而从模型层面学习多层级动态感知能力。在涵盖静态属性、动态物理特征及多层级实际操作任务的基准上评估,结果表明其在不同传感器与任务上均表现一致且优越。
原文摘要 · Abstract (English)
Real-world contact-rich manipulation demands robots to perceive temporal tactile feedback, capture subtle surface deformations, and reason about object properties as well as force dynamics. Although optical tactile sensors are uniquely capable of providing such rich information, existing tactile datasets and models remain limited. These resources primarily focus on object-level attributes (e.g., material) while largely overlooking fine-grained tactile temporal dynamics during physical interactions. We consider that advancing dynamic tactile perception requires a systematic hierarchy of dynamic perception capabilities to guide both data collection and model design. To address the lack of tactile data with rich dynamic information, we present ToucHD, a large-scale hierarchical tactile dataset spanning tactile atomic actions, real-world manipulations, and touch-force paired data. Beyond scale, ToucHD establishes a comprehensive tactile dynamic data ecosystem that explicitly supports hierarchical perception capabilities from the data perspective. Building on it, we propose AnyTouch 2, a general tactile representation learning framework for diverse optical tactile sensors that unifies object-level understanding with fine-grained, force-aware dynamic perception. The framework captures both pixel-level and action-specific deformations across frames, while explicitly modeling physical force dynamics, thereby learning multi-level dynamic perception capabilities from the model perspective. We evaluate our model on benchmarks that covers static object properties and dynamic physical attributes, as well as real-world manipulation tasks spanning multiple tiers of dynamic perception capabilities-from basic object-level understanding to force-aware dexterous manipulation. Experimental results demonstrate consistent and strong performance across sensors and tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。