通过分析触觉运动相关性,提升接触操作中触觉细节的识别能力。
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

- 利用瞬时与累积运动的相关性建模触觉动态,增强细粒度接触状态区分能力。
- 在Dexterity Benchmark上实现92.3%的接触分类准确率,显著优于基线方法。
- 适合需要高精度触觉感知的机器人抓取与装配任务研究者。
基于光学触觉传感器的视觉-触觉策略在高接触密度操作中展现出巨大潜力。这类传感器通过内部相机监测弹性凝胶表面的形变,实现高空间分辨率和多维力感知,间接推断触觉信息。然而,提取接触密集操作所需的精细接触状态仍是开放挑战。现有方法通常使用原始图像或累积运动场表示触觉线索,但二者均易产生感知模糊:原始图像主要捕捉外观变化,累积运动场仅反映整体形变。因此,不同的精细接触状态可能呈现高度相似的模式,难以明确区分细微差异。为此,我们探索触觉运动的动态先验,发现瞬时与累积运动的相关性可有效区分精细接触状态。基于此,提出一种运动感知的触觉表征以促进接触密集操作。此外,视觉与触觉模态的有效融合同样关键。多数现有融合方法或直接拼接各模态特征,或分别训练模态专用网络再融合输出,难以同时建模跨模态交互并保留模态特异性。本文采用Mixture-of-Transformers架构,提出统一的模态感知视觉-触觉策略,既捕捉跨模态互补性,又保持模态特有属性。
原文摘要 · Abstract (English)
Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial resolution and multi-dimensional force sensing by utilizing an internal camera to monitor the deformation of their elastic gel surface, thereby indirectly inferring tactile cues. Despite their advantages, extracting fine-grained contact states necessary for contact-rich manipulation remains an open challenge. Existing methods typically use either raw images or cumulative motion fields to represent tactile cues. However, both are prone to perception ambiguity. Raw tactile images mainly capture appearance changes, while cumulative motion fields only reflect the aggregate gel deformation. Consequently, distinct fine-grained contact states can exhibit highly similar patterns, making it difficult to explicitly distinguish subtle contact variations. To address this issue, we explore the dynamic priors of tactile motion and discover that the correlation between transient and cumulative motion can explicitly distinguish fine-grained contact states. Based on this insight, we propose a motion-aware tactile representation to facilitate contact-rich manipulation. Beyond tactile representation, effective fusion of tactile and visual modalities is also critical. Most existing fusion methods either directly concatenate features from each modality or train modality-specific networks separately and fuse their outputs. However, these strategies struggle to simultaneously model cross-modal interactions and preserve modality-specific characteristics. In this work, we take advantage of the Mixture-of-Transformers architecture and propose a unified modality-aware visuo-tactile policy that captures cross-modal complementarity while maintaining modality-specific properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。