arXiv:2512.01908cs.CV2025-12

提出空间感知自监督学习,提升机器人触觉视觉融合感知能力

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

  • 在BYOL基础上引入三类地图级损失,保留特征空间结构
  • 边缘位姿回归任务误差仅0.3955,比次优方法降低30%
  • 适合需要精细几何感知的机器人操控任务

接触丰富的机器人操作需要编码局部几何特性的表征。视觉提供全局上下文但缺乏纹理、硬度等直接测量,而触觉可补充这些线索。现代多模态传感器将视觉与触觉融合为单一对齐图像,特别适合需要双模态信息的操作任务。然而多数自监督学习框架将特征图压缩为全局向量,丢失空间结构,与操作需求不符。为此,我们提出SARL,一种空间感知自监督框架,在Bootstrap Your Own Latent(BYOL)架构中加入三类地图级目标:显著性对齐(SAL)、块原型分布对齐(PPDA)和区域亲和匹配(RAM),以保持关注点、部件组合与几何关系在不同视图间的一致性。这些损失作用于中间特征图,补充全局目标。SARL在六项下游任务中均优于九种自监督基线方法。在对几何敏感的边缘位姿回归任务中,其平均绝对误差(MAE)达0.3955,相较次优自监督方法(0.5682)相对提升30%,接近监督上限。结果表明,对于融合的视觉-触觉数据,最有效的信号是结构化的空间等变性,即特征随物体几何规律变化,从而实现更强大的机器人感知。

原文摘要 · Abstract (English)

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these cues. Modern visuo-tactile sensors capture both modalities in a single fused image, yielding intrinsically aligned inputs that are well suited to manipulation tasks requiring visual and tactile information. Most self-supervised learning (SSL) frameworks, however, compress feature maps into a global vector, discarding spatial structure and misaligning with the needs of manipulation. To address this, we propose SARL, a spatially-aware SSL framework that augments the Bootstrap Your Own Latent (BYOL) architecture with three map-level objectives, including Saliency Alignment (SAL), Patch-Prototype Distribution Alignment (PPDA), and Region Affinity Matching (RAM), to keep attentional focus, part composition, and geometric relations consistent across views. These losses act on intermediate feature maps, complementing the global objective. SARL consistently outperforms nine SSL baselines across six downstream tasks with fused visual-tactile data. On the geometry-sensitive edge-pose regression task, SARL achieves a Mean Absolute Error (MAE) of 0.3955, a 30% relative improvement over the next-best SSL method (0.5682 MAE) and approaching the supervised upper bound. These findings indicate that, for fused visual-tactile data, the most effective signal is structured spatial equivariance, in which features vary predictably with object geometry, which enables more capable robotic perception.

自监督学习触觉感知机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。