arXiv:2604.01414cs.RO2026-04被引 3

让机器人在接触时用力矩,不接触时不用,提升操作成功率14%。

Learning When to See and When to Feel: Adaptive Vision-Torque Fusion for Contact-Aware Manipulation

论文配图:Learning When to See and When to Feel: Adaptive Vision-Torque Fusion for Contact-Aware Manipulation
图 1 · 摘自论文原文
  • 根据是否接触动态切换视觉与力矩信息的使用
  • 在接触任务中成功率比最强基线高14%
  • 适合需要精准力控的抓取、插入等场景

基于视觉的策略在机器人操作中表现良好,得益于视觉观测的易获取性和丰富性。但在接触密集且对力敏感的任务中,仅靠视觉难以捕捉接触动态、对齐状态和交互质量,力/扭矩(F/T)信号至关重要。尽管已有多种融合视觉与F/T信号的方法,如辅助预测目标、专家混合架构和接触感知门控机制,但这些方法缺乏系统比较。本文在基于扩散模型的操作策略中,对不同F/T-视觉融合策略进行受控对比。此外,提出一种自适应融合策略:在非接触阶段忽略F/T信号,在接触阶段则自适应地结合视觉与扭矩信息。实验表明,该方法在成功率上优于最强基线14%,凸显了接触感知多模态融合在机器人操作中的重要性。

原文摘要 · Abstract (English)

Vision-based policies have achieved a good performance in robotic manipulation due to the accessibility and richness of visual observations. However, purely visual sensing becomes insufficient in contact-rich and force-sensitive tasks where force/torque (F/T) signals provide critical information about contact dynamics, alignment, and interaction quality. Although various strategies have been proposed to integrate vision and F/T signals, including auxiliary prediction objectives, mixture-of-experts architectures, and contact-aware gating mechanisms, a comparison of these approaches remains lacking. In this work, we provide a controlled comparison of different F/T-vision integration strategies within diffusion-based manipulation policies. In addition, we propose an adaptive integration strategy that ignores F/T signals during non-contact phases while adaptively leveraging both vision and torque information during contact. Experimental results demonstrate that our method outperforms the strongest baseline by 14% in success rate, highlighting the importance of contact-aware multimodal fusion for robotic manipulation.

多模态融合力控操作扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。