arXiv:2604.11579cs.CV2026-04被引 1

通过触觉引导视觉定位材料区域,提升细粒度匹配能力

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

论文配图:Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions
图 1 · 摘自论文原文
  • 基于密集跨模态特征交互学习局部触觉-视觉对齐
  • 在新数据集上触觉定位准确率显著超越已有方法
  • 适合研究多模态感知与真实场景材料识别的学者

我们解决触觉定位问题,目标是识别图像中与触觉输入具有相同材质属性的区域。现有视觉触觉方法依赖全局对齐,难以捕捉该任务所需的细粒度局部对应关系。这一挑战因现有数据集主要包含近距离、低多样性图像而加剧。本文提出一种模型,通过密集跨模态特征交互学习局部视觉-触觉对齐,生成触觉显著性图,实现触觉条件下的材质分割。为克服数据集局限,我们引入:(i) 原生多材质场景图像以扩展视觉多样性;(ii) 材质多样性配对策略,将每个触觉样本与视觉多样但触觉一致的图像配对,增强上下文定位能力并提高对弱信号的鲁棒性。我们还构建了两个新的触觉锚定材质分割数据集用于量化评估。在新旧基准上的实验表明,该方法在触觉定位任务中显著优于以往视觉触觉方法。

原文摘要 · Abstract (English)

We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The challenge is amplified by existing datasets, which predominantly contain close-up, low-diversity images. We propose a model that learns local visuo-tactile alignment via dense cross-modal feature interactions, producing tactile saliency maps for touch-conditioned material segmentation. To overcome dataset constraints, we introduce: (i) in-the-wild multi-material scene images that expand visual diversity, and (ii) a material-diversity pairing strategy that aligns each tactile sample with visually varied yet tactilely consistent images, improving contextual localization and robustness to weak signals. We also construct two new tactile-grounded material segmentation datasets for quantitative evaluation. Experiments on both new and existing benchmarks show that our approach substantially outperforms prior visuo-tactile methods in tactile localization.

多模态感知触觉定位材质分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。