arXiv:2511.19509cs.LG2025-11AAAI被引 7

用自适应融合提升触觉视觉多模态材料识别鲁棒性

TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception

  • 通过自适应门控和跨模态注意力动态融合多模态信息
  • 在SSMC和USMC任务上分别提升2.48%和6.83%准确率
  • 适合机器人环境感知与高安全场景应用

基于视觉的材料感知方法在视觉受限条件下性能显著下降,推动非视觉多模态感知的发展。然而现有方法常采用简单融合策略,忽视模态噪声、模态缺失及模态重要性动态变化等挑战,导致基准任务表现不佳。本文提出稳健的多模态融合框架TouchFormer,引入模态自适应门控(MAG)机制与模态内/间注意力机制,实现跨模态特征的自适应集成。此外,设计跨实例嵌入正则化(CER)策略,在细粒度子类材料识别任务中显著提升准确率。实验表明,相比现有非视觉方法,TouchFormer在SSMC和USMC任务上分别提升2.48%和6.83%。真实机器人实验验证其在环境感知中的有效性,为应急响应与工业自动化等安全关键场景部署提供可能。代码与数据集将开源,视频见附录。

原文摘要 · Abstract (English)

Traditional vision-based material perception methods often experience substantial performance degradation under visually impaired conditions, thereby motivating the shift toward non-visual multimodal material perception. Despite this, existing approaches frequently perform naive fusion of multimodal inputs, overlooking key challenges such as modality-specific noise, missing modalities common in real-world scenarios, and the dynamically varying importance of each modality depending on the task. These limitations lead to suboptimal performance across several benchmark tasks. In this paper, we propose a robust multimodal fusion framework, TouchFormer. Specifically, we employ a Modality-Adaptive Gating (MAG) mechanism and intra- and inter-modality attention mechanisms to adaptively integrate cross-modal features, enhancing model robustness. Additionally, we introduce a Cross-Instance Embedding Regularization(CER) strategy, which significantly improves classification accuracy in fine-grained subcategory material recognition tasks. Experimental results demonstrate that, compared to existing non-visual methods, the proposed TouchFormer framework achieves classification accuracy improvements of 2.48% and 6.83% on SSMC and USMC tasks, respectively. Furthermore, real-world robotic experiments validate TouchFormer's effectiveness in enabling robots to better perceive and interpret their environment, paving the way for its deployment in safety-critical applications such as emergency response and industrial automation. The code and datasets will be open-source, and the videos are available in the supplementary materials.

多模态感知机器人材料识别自适应融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。