arXiv:2508.14441cs.RO2025-08被引 5

用触觉与视觉动态融合提升机器人灵巧操作能力

FBI: Learning Dexterous In-hand Manipulation with Dynamic Visuotactile Shortcut Policy

  • 通过动态融合触觉与视觉信息,建立触觉信号与物体运动的因果联系
  • 在仿真和真实场景中,于多个任务上超越基线方法性能
  • 适合研究多模态感知、灵巧操作或机器人模仿学习的开发者

灵巧的物体抓握操作是机器人领域长期存在的挑战,源于复杂的接触动力学和观测不完全性。尽管人类能协调视觉与触觉完成此类任务,但现有机器人方法常偏重单一模态,限制了适应性。本文提出流动先于模仿(FBI)框架,一种基于视觉-触觉仿人学习的方法,通过运动动力学动态融合触觉交互与视觉观测。不同于以往静态融合方式,FBI利用动力学感知潜在模型建立触觉信号与物体运动间的因果关系。该方法采用基于Transformer的交互模块,融合由运动流提取的触觉特征与视觉输入,训练一步扩散策略以实现实时执行。大量实验表明,该方法在两个自定义抓握任务及三个标准灵巧操作任务中,均在仿真与真实世界中优于基线方法。代码、模型与更多结果见:https://sites.google.com/view/dex-fbi。

原文摘要 · Abstract (English)

Dexterous in-hand manipulation is a long-standing challenge in robotics due to complex contact dynamics and partial observability. While humans synergize vision and touch for such tasks, robotic approaches often prioritize one modality, therefore limiting adaptability. This paper introduces Flow Before Imitation (FBI), a visuotactile imitation learning framework that dynamically fuses tactile interactions with visual observations through motion dynamics. Unlike prior static fusion methods, FBI establishes a causal link between tactile signals and object motion via a dynamics-aware latent model. FBI employs a transformer-based interaction module to fuse flow-derived tactile features with visual inputs, training a one-step diffusion policy for real-time execution. Extensive experiments demonstrate that the proposed method outperforms the baseline methods in both simulation and the real world on two customized in-hand manipulation tasks and three standard dexterous manipulation tasks. Code, models, and more results are available in the website https://sites.google.com/view/dex-fbi.

灵巧操作多模态融合模仿学习触觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。