用最优传输机制融合视觉、力觉与位姿,提升复杂操作的机器人学习成功率。
Spacetime Optimal-Transport Attention for Visuo-Haptic Imitation Learning of Contact-Rich Manipulation

- 基于最优传输构建三模态注意力,显式约束接触区域的选择
- 真实机器人实验中成功率100%,抗干扰能力显著优于传统方法
- 适合需要高精度触觉交互的工业装配与表面处理任务
接触密集型操作如精密插装、连接器对接、抛光和贴合擦拭等,因耦合了不连续接触动力学、观测不全及严格安全约束,难以由数据驱动控制器完成。单一传感模态不足:视觉提供接触前全局信息,力/扭矩(F/T)反馈控制接触后交互,本体感知位姿则提供一致的运动学基础。以往模仿学习策略多仅使用单模态或双模态信号,少数三模态融合方法采用无先验的通用注意力模块,未考虑注意力分布的任务相关性。本文提出时空最优传输注意力(SO-TA),以熵正则化最优传输(OT)取代软最大归一化的块注意力,将力-位姿生成的子查询与视觉块进行对齐。显式的边际约束作为结构化归纳偏置,促进对接触区域的条件感知空间选择,在光照变化、干扰物和部分遮挡下保持稳定。SO-TA与基于扩散模型的序列策略结合,将观察窗口映射为姿态-动作片段。在三个真实机器人任务上评估:精密插销装配、BCM线缆连接器插入、曲面标记擦除。每条件下约200次采样,SO-TA在精密插销任务上达到100%成功率,相较匹配容量的交叉注意力提升至93%;在光照、干扰物和部分遮挡扰动下仍保持82.5%成功率,而拼接基线降至43.5%。OT生成的块热图与留一法模态影响比提供可解释、阶段依赖的诊断信息。
原文摘要 · Abstract (English)
Contact-rich manipulation tasks such as tight-clearance insertion, connector mating, polishing, and surface-conforming wiping remain difficult for data-driven controllers because they couple discontinuous contact dynamics, partial observability, and strict safety constraints. No single sensing modality suffices: vision supplies global context before contact, force/torque (F/T) feedback governs interaction after contact, and proprioceptive pose provides a consistent kinematic backbone. Most prior imitation-learning policies for contact-rich tasks operate on uni- or bi-modal signals, and the few that fuse three modalities typically adopt off-the-shelf attention modules with no explicit prior on how attention mass should be distributed across task-relevant regions. We present Spacetime Optimal-Transport Attention (SO-TA), a tri-modal fusion backbone that replaces softmax-normalized patch attention by an entropy-regularized Optimal Transport (OT) alignment between force-pose-derived sub-queries and visual patches. Explicit marginal constraints act as a structured inductive bias for contact-rich tasks, encouraging conditioning-aware spatial selection that is stable across illumination, distractors, and partial occlusion. SO-TA is paired with a diffusion-based sequence policy mapping observation windows to pose-action chunks. We evaluate SO-TA on three real-robot tasks: tight peg-in-hole assembly, BCM wiring-connector insertion, and curved-surface mark erasing. With ~200 rollouts per condition, SO-TA reaches 100% success on tight peg-in-hole versus 93% for cross-attention at matched capacity, and retains 82.5% success under illumination, distractor, and partial-occlusion perturbations where a concatenation baseline drops to 43.5%. OT-derived patch heatmaps and leave-one-out modality-influence ratios provide interpretable, phase-dependent diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。