让机器人通过触觉预测学习更精准的物理接触动作。
TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

- 用多模态融合编码触觉信息,保留力与形变的全局特征。
- 触觉历史建模使未来预测反映力和形变的动态变化。
- 分离触觉与动作信号,确保触觉监督不干扰动作决策。
世界动作模型(WAMs)结合未来状态预测与机器人动作生成,但现有方法主要依赖视觉未来。视觉预测虽能捕捉场景结构和物体运动,却难以提供接触操作中力、形变、剪切和滑移的充分监督。为此,我们提出机械感知触觉WAM(TacWAM),分三步解决该问题:首先,空间对齐融合(SAF)触觉编码器将触觉外观、密集力场和形变流映射至共享潜在预测空间,双侧力与扭矩重建保持全局接触信息;其次,触觉历史编码器提供时间上下文,使未来触觉预测反映力与形变的动态演化;第三,锚点引导三模态(AGT)注意力分离当前视觉/触觉锚点、未来预测项与动作项,使未来触觉状态可监督训练,但不被动作分支直接读取。我们在四个真实世界接触密集型操作任务上评估:脆弱抓取、持续表面接触及动态手中操纵。TacWAM平均成功率75.0%,比最强基线高出37.5个百分点。分阶段消融实验显示,移除触觉历史或放松未来预测目标均导致性能持续下降。结果表明,结合信息丰富的触觉表示与部署一致的信息约束,未来触觉监督可有效提升接触感知动作学习。
原文摘要 · Abstract (English)
World Action Models (WAMs) combine future-state prediction with robot action generation, but existing approaches largely rely on visual futures. Visual prediction captures scene structure and object motion, yet provides limited supervision for force, deformation, shear, and slip during contact-rich manipulation. This creates two design requirements: tactile futures should carry meaningful physical information, and they should not become privileged cues for action generation. We present TacWAM, a mechanics-aware tactile WAM that addresses this challenge in three steps. First, a Spatially Aligned Fusion (SAF) Tactile Encoder maps tactile appearance, dense force fields, and deformation flow into a shared latent prediction space, with bilateral force and torque reconstruction preserving global contact information. Second, a tactile history encoder provides temporal context so future tactile prediction reflects how force and deformation change beyond the current tactile observation. Third, Anchor-Guided Tri-Modal (AGT) Attention separates current visual and tactile anchors, future prediction tokens, and action tokens, allowing future tactile states to supervise training without being directly read by the action branch. We evaluate TacWAM on four real-world contact-rich manipulation tasks covering fragile grasping, sustained surface contact, and dynamic in-hand manipulation. TacWAM achieves an average success rate of 75.0%, exceeding the strongest evaluated baseline by 37.5 percentage points. Staged ablations show consistent degradation when tactile history is removed and access to future prediction targets is relaxed. These results indicate that future tactile supervision can improve contact-aware action learning when combined with informative tactile representations and deployment-consistent information constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。