arXiv:2603.05687cs.RO2026-03被引 3

让机器人通过触觉预测接触状态,实现更灵巧的抓取与操作。

Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding

  • 用扩散模型预测触觉与机器人状态的协同变化轨迹。
  • 在真实四指手和模拟五指手上,任务成功率显著优于基线方法。
  • 适合研究灵巧操作、触觉感知或具身智能的科研人员。

多指灵巧操作中,任务成功依赖于动态演变的多点接触,对物体几何、摩擦变化和滑移极为敏感。尽管触觉引导策略已有进展,但多数仅将触觉信号作为附加观测,未建模接触状态或动作输出与底层控制器动力学的交互。本文提出接触接地策略(CGP),一种视觉-触觉联合策略,通过预测实际机器人状态与触觉反馈的耦合轨迹,并利用学习的接触一致性映射,将预测结果转化为可执行的目标状态,供柔顺控制器使用。CGP由两部分组成:(i) 条件扩散模型,在压缩隐空间中预测未来机器人状态与触觉反馈;(ii) 学习的接触一致性映射,将预测的机器人状态-触觉对转换为柔顺控制器可执行的目标。我们在搭载Digit360指尖触觉传感器的真实四指Allegro V5机械手以及具有密集全手触觉阵列的模拟五指Tesollo DG-5F机械手上评估了CGP。在翻转、精细抓握和工具使用等多样灵巧操作任务中,CGP均优于视觉-运动和视觉-触觉扩散策略基线。

原文摘要 · Abstract (English)

Contact-rich dexterous manipulation with multi-finger hands remains an open challenge in robotics because task success depends on multi-point contacts that continuously evolve and are highly sensitive to object geometry, frictional transitions, and slip. Recently, tactile-informed manipulation policies have shown promise. However, most use tactile signals as additional observations rather than modeling contact state or how their action outputs interact with low-level controller dynamics. We present Contact-Grounded Policy (CGP), a visuotactile policy that grounds multi-point contacts by predicting coupled trajectories of actual robot state and tactile feedback, and using a learned contact-consistency mapping to convert these predictions into executable target robot states for a compliance controller. CGP consists of two components: (i) a conditional diffusion model that forecasts future robot state and tactile feedback in a compressed latent space, and (ii) a learned contact-consistency mapping that converts the predicted robot state-tactile pair into executable targets for a compliance controller, enabling it to realize the intended contacts. We evaluate CGP using a physical four-finger Allegro V5 hand with Digit360 fingertip tactile sensors, and a simulated five-finger Tesollo DG-5F hand with dense whole-hand tactile arrays. Across a range of dexterous tasks including in-hand manipulation, delicate grasping, and tool use, CGP outperforms visuomotor and visuotactile diffusion-policy baselines.

灵巧操作触觉感知扩散模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。