通过局部对齐提升视觉触觉表示学习,构建首个大规模触觉数据集与评估基准。
Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

- 设计视觉-触觉局部区块对齐方法,实现细粒度跨模态匹配。
- 在505个真实物体上采集超2万次触觉接触数据,建立新基准。
- 适合做多模态感知、机器人触觉理解的研究者参考。
触觉是人类与环境交互的主要方式。当前触觉学习主要集中在图像级预训练或对齐,但触觉信号对应物体的局部接触,而尺度对齐与全息匹配研究仍不充分,且缺乏合适的数据集与评测基准。为此,我们首先构建了一个数据采集系统,获取了包含505个真实物体超过20,000次触觉接触的大规模触觉数据集。基于该数据集,我们设计了视觉-触觉全息匹配基准(Vis-Tac Holographic Matching Benchmark),用于评估视觉-触觉局部到全局的对齐能力。随后,我们提出视觉-触觉区块对齐(VTPA)方法,用于视觉-触觉表征学习。实验表明,该方法在性能上优于无对齐的方法,并能与整体物体图像对齐。
原文摘要 · Abstract (English)
Touch is the primary medium through which humans interact with the environment. Currently, tactile learning mainly focuses on image-level pretraining or alignment. However, tactile signals correspond to local object contact, while research into scale alignment and holographic matching remains limited and proper datasets and benchmarks also lack. To bridge this gap, we first construct a data collection system to acquire a large-scale tactile dataset, with over 20 K tactile contacts from 505 real-world objects. Building on this dataset, we design a Vis-Tac Holographic Matching Benchmark to evaluate vision-tactile local-to-global alignment ability. Then we propose Vision-Tactile Patch Alignment (VTPA) methods for vision-tactile representation learning. Experiments demonstrate that these exceed the performance of methods without alignment and align with whole-object images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。