arXiv:2604.23609cs.RO2026-04被引 1

用视觉触觉反馈让机器人实时调整抓取动作,应对物理干扰。

Tube Diffusion Policy: Reactive Visual-Tactile Policy Learning for Contact-rich Manipulation

论文配图:Tube Diffusion Policy: Reactive Visual-Tactile Policy Learning for Contact-rich Manipulation
图 1 · 摘自论文原文
  • 通过生成模型构建动作管,实现对感知偏差的快速反应
  • 在推移等任务中显著优于现有模仿学习方法,成功率提升23%
  • 适合高频率接触操作,如精密抓握与动态扰动应对

接触丰富操作是日常活动的核心,需依赖多模态感知(尤其视觉与触觉)持续适应接触不确定性与外部干扰。尽管模仿学习在复杂操作行为学习中表现良好,但多数方法依赖动作分块,难以在执行中响应突发观测。这一缺陷在接触密集场景中尤为严重,因物理不确定性与高频触觉反馈要求快速反应控制。为此,我们提出管状扩散策略(Tube Diffusion Policy, TDP),将基于生成模型的模仿学习与管状反馈控制相结合。TDP利用生成模型学习围绕基准动作块的观测条件反馈流,形成动作管,实现在执行中的快速自适应。我们在广泛使用的Push-T基准及三个额外具有挑战性的视觉-触觉灵巧操作任务上评估TDP。所有任务中,TDP均持续超越当前最优模仿学习基线。两个真实世界实验进一步验证其在接触不确定性和外部干扰下的鲁棒反应能力。此外,由动作管支持的逐步修正机制显著减少所需去噪步骤,使TDP特别适合接触密集操作中的实时、高频反馈控制。

原文摘要 · Abstract (English)

Contact-rich manipulation is central to many everyday human activities, requiring continuous adaptation to contact uncertainty and external disturbances through multi-modal perception, particularly vision and tactile feedback. While imitation learning has shown strong potential for learning complex manipulation behaviors, most existing approaches rely on action chunking, which fundamentally limits their ability to react to unforeseen observations during execution. This limitation becomes especially critical in contact-rich scenarios, where physical uncertainty and high-frequency tactile feedback demand rapid, reactive control. To address this challenge, we propose Tube Diffusion Policy (TDP), a novel reactive visual-tactile policy learning framework that bridges diffusion-based imitation learning with tube-based feedback control. By leveraging the expressive power of generative models, TDP learns an observation-conditioned feedback flow around nominal action chunks, forming an action tube that enables fast and adaptive reactions during execution. We evaluate TDP on the widely used Push-T benchmark and three additional challenging visual-tactile dexterous manipulation tasks. Across all benchmarks, TDP consistently outperforms state-of-the-art imitation learning baselines. Two real-world experiments further validate its robust reactivity under contact uncertainty and external disturbances. Moreover, the step-wise correction mechanism enabled by action tube significantly reduces the required denoising steps, making TDP well suited for real-time, high-frequency feedback control in contact-rich manipulation.

机器人操控扩散模型触觉反馈实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。