用同一模型同时学会找孔和精准插入,提升机器人装配鲁棒性。
SI-Diff: A Framework for Learning Search and High-Precision Insertion with a Force-Domain Diffusion Policy

- 基于力域扩散模型统一建模搜索与高精度插入行为。
- 在5毫米偏移下仍可成功装配,优于现有方法2毫米极限。
- 无需换模型即可零样本迁移至未知形状,适合复杂装配场景。
接触密集型装配是机器人领域的基础任务,但因相对位姿不确定性(如插孔任务中的错位与微小间隙)面临巨大挑战。现有方法通常将搜索与高精度插入分开处理,因二者动作模式差异大。然而,希望在一个模型中同时支持两项任务,避免模型切换。本文提出SI-Diff框架,通过力域扩散策略联合学习搜索与高精度插入。为此,引入新式模态条件机制,使单一框架能捕捉不同动作行为;同时设计新的搜索教师策略,生成多样化轨迹。通过在教师策略提供的高效成功示范上训练,模型学习从触觉与末端执行器速度观测到有效动作的映射。实验表明,相较于最先进基线TacDiffusion,SI-Diff将x-y方向容错范围从2毫米扩展至5毫米,并展现出对未见形状的强大零样本迁移能力。
原文摘要 · Abstract (English)
Contact-rich assembly is fundamental in robotics but poses significant challenges due to uncertainties in relative poses, such as misalignments and small clearances in peg-in-hole tasks. Existing approaches typically address search and high-precision insertion separately, because these tasks involve distinct action patterns. However, supporting both tasks within a single model, without switching models or weights, is desirable for intelligent assembly systems. In this work, we propose SI-Diff, a framework that learns both search and high-precision insertion through a force-domain diffusion policy. To this end, we introduce a new mode-conditioning mechanism that enables the policy to capture distinct action behaviors under a single framework. Moreover, we develop a new search teacher policy that can generate diverse trajectories. By training on successful and efficient demonstrations provided by the teacher policy, the model learns the mapping from tactile and end-effector velocity observations to effective action behaviors. We conduct thorough experiments to show that SI-Diff extends the tolerance to x-y misalignments from 2 mm to 5 mm compared to the state-of-the-art baseline, TacDiffusion, while also demonstrating strong zero-shot transferability to unseen shapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。