视觉与力觉融合的端到端控制,提升复杂操作的反应速度和成功率。
ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning
- 用因果注意力同步处理视觉慢信号和力觉快信号
- 在接触任务中成功率提升23%,反应速度更快
- 适合需要高精度触觉反馈的机器人操作场景
人类级的高接触操作依赖于两种关键模态:视觉提供空间丰富但时间滞后的全局上下文,力觉感知捕捉快速的局部接触动态。由于二者在频率和信息量上的根本差异,融合极具挑战。本文提出ImplicitRDP,一种统一的端到端视觉-力觉扩散策略,将视觉规划与力觉反应控制整合于单一网络。引入结构化慢-快学习机制,利用因果注意力同时处理异步的视觉与力觉令牌,使策略能在动作频率下实现快速力控,同时保持动作块的时间一致性。此外,为缓解端到端模型中模态权重失衡问题,提出基于虚拟目标的表征正则化,将力反馈映射至与动作相同的语义空间,提供比原始力预测更强的物理驱动学习信号。大量接触密集型任务实验表明,ImplicitRDP显著优于仅视觉或分层基线,以更简化的训练流程实现更高反应性与成功率。代码与视频见https://implicit-rdp.github.io。
原文摘要 · Abstract (English)
Human-level contact-rich manipulation relies on the distinct roles of two key modalities: vision provides spatially rich but temporally slow global context, while force sensing captures rapid local contact dynamics. Integrating these signals is challenging due to their fundamental frequency and informational disparities. In this work, we propose ImplicitRDP, a unified end-to-end visual-force diffusion policy that integrates visual planning and reactive force control within a single network. We introduce Structural Slow-Fast Learning, a mechanism utilizing causal attention to simultaneously process asynchronous visual and force tokens, allowing the policy to perform rapid force control at the action rate while maintaining the temporal coherence of action chunks. Furthermore, to mitigate modality collapse where end-to-end models fail to adjust the weights across different modalities, we propose Virtual-target-based Representation Regularization. This auxiliary objective maps force feedback into the same space as the action, providing a stronger, physics-grounded learning signal than raw force prediction. Extensive experiments on contact-rich tasks demonstrate that ImplicitRDP significantly outperforms both vision-only and hierarchical baselines, achieving superior reactivity and success rates with a streamlined training pipeline. Code and videos are available at https://implicit-rdp.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。