arXiv:2608.01102cs.RO2026-08

通过触觉动态掩码与注意力增强,提升触觉丰富的操作任务学习效率。

CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation

论文配图:CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation
图 1 · 摘自论文原文
  • 基于接触状态动态调节视觉与触觉信息的注意力权重
  • 在仿真中使成功率提升18个百分点,真实场景达60%成功率
  • 无需修改原有模型,适配多种策略架构,适合机器人操控研究

在接触丰富的操作任务中,视觉信息主导自由空间运动,而触觉信息在接触时更具价值。然而,传统的基于Transformer的视听融合策略通常采用拼接或可学习门控,缺乏显式的接触先验,难以从示范数据中高效学习跨模态表征。为此,我们提出轻量级的CAAT框架,通过注意力缩放和动态触觉掩码显式引入接触先验。CAAT在接触前强调视觉信息,在接触时增强触觉信息,并通过对比当前触觉观测与无接触参考,抑制静态背景信息。该方法可无缝集成至常用Transformer策略中,无需修改动作解码器。仿真结果表明,结合CAAT的ACT比直接融合提升18.0个百分点,比门控融合提升10.0个百分点。在真实世界中使用视觉-触觉UMI平台,CAAT在ACT、Diffusion Policy和$π_0$上平均成功率达60.0%,优于最强基线平均21.1个百分点。结果证明,显式接触先验与动态触觉掩码能有效提升视听策略学习与任务性能。

原文摘要 · Abstract (English)

In contact-rich manipulation, visual observations primarily guide motion in free space, whereas tactile observations become particularly informative during contact. However, standard Transformer-based visuo-tactile policies typically rely on either token concatenation or learnable gating. These approaches lack explicit contact-aware priors, making it difficult to efficiently learn effective cross-modal representations from demonstrations. To address this limitation, we propose CAAT, a lightweight contact-aware framework that explicitly incorporates contact priors through attention scaling and dynamic tactile masking. Specifically, CAAT emphasizes visual information before contact and tactile information during contact. It also suppresses static background tokens by comparing the current tactile observation with a non-contact reference. CAAT can be integrated into commonly used Transformer-based policies without modifying their action decoders. In simulation, integrating CAAT with ACT improves the average success rate by 18.0 percentage points over direct visuo-tactile fusion and by 10.0 percentage points over gated fusion. In real-world experiments using a visuo-tactile UMI platform, CAAT achieves an average success rate of 60.0% across ACT, Diffusion Policy, and $π_0$, outperforming the strongest baseline by an average of 21.1 percentage points. These results demonstrate that explicit contact priors and dynamic tactile masking are effective in improving visuo-tactile policy learning and task performance of diverse policy architectures. https://mrjiangjm.github.io/caat/

机器人操控触觉感知Transformer多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。