arXiv:2606.19161cs.RO2026-06被引 3

用视觉+全手触觉数据构建大模型基准,提升机器人触觉表征能力。

HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision

论文配图:HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision
图 1 · 摘自论文原文
  • 结合第一视角视觉与全手触觉数据,构建多任务评估框架。
  • 在4项任务中超越基线,触觉相似度检索召回率提升至85.23%。
  • 适合研究机器人触觉学习、多模态感知的科研人员。

由于触觉传感器设计、数据格式和机器人本体的多样性,建立通用的机器人操作触觉表征学习基准仍具挑战。本文提出一种可扩展的未来方向:结合第一视角视觉与全手触觉数据。为此,我们引入 extbf{HT-Bench},一个大规模多任务基准,包含1000万帧RGB图像和780万帧触觉数据,覆盖226个任务。该基准从三方面评估触觉表征:是否编码有意义的接触几何、能否对齐触觉与视觉信息、是否能泛化至未见任务。涵盖四项任务:细粒度触觉相似性检索、掩码触觉补全、视觉到触觉合成、多模态触觉帧预测。我们进一步提出 extbf{HandTouch},一种通过渐进式空间、跨模态和时间训练的向量量化视觉-触觉编码器。在HT-Bench上,HandTouch持续优于代表性触觉编码器基线:细粒度触觉相似性检索的Recall@5从74.65%提升至85.23%,掩码触觉补全的RMSE从0.022降至0.010,视觉到触觉合成的OOD cIoU从0.628升至0.705。结果表明,大规模第一视角全手触觉数据为精细操作中的触觉表征学习提供了可扩展评估基础。

原文摘要 · Abstract (English)

Establishing a universal benchmark for tactile representation learning in robotic manipulation remains challenging due to the diversity of tactile sensor designs, data formats, and robot embodiments. Rather than seeking to establish such, we explore a scalable and promising direction for future development: egocentric vision paired with full-hand tactile data. To this end, we introduce \textbf{HT-Bench}, a large-scale multi-task benchmark for dexterous full-hand tactile sensing, comprising 10M RGB frames and 7.8M tactile frames collected across 226 tasks. HT-Bench evaluates tactile representations from three key perspectives: whether they encode meaningful contact geometry, whether they can align tactile observations with visual information, and whether they generalize to unseen tasks. To assess these capabilities, HT-Bench includes four tasks: fine-grained tactile similarity retrieval, masked tactile inpainting, vision-to-tactile synthesis, and multimodal tactile frame prediction. We further propose \textbf{HandTouch}, a vector-quantized vision--tactile encoder that learns tactile representations through progressive spatial, cross-modal, and temporal training. Across HT-Bench, HandTouch consistently outperforms representative tactile encoder baselines, improving Recall@5 on fine-grained tactile similarity retrieval from 74.65\% to 85.23\%, reducing RMSE on masked tactile inpainting from 0.022 to 0.010, and increasing OOD cIoU on vision-to-tactile synthesis from 0.628 to 0.705. These results demonstrate the effectiveness of HandTouch and suggest that large-scale egocentric full-hand tactile data provides a scalable basis for evaluating and advancing tactile representation learning in dexterous manipulation.

触觉表征多模态学习机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。