arXiv:2602.23253cs.RO2026-02被引 1

用仿真+现实残差策略,让机器人装配更准更快。

SPARR: Simulation-based Policies with Asymmetric Real-world Residuals for Assembly

  • 先用仿真训练基础策略,再用现实数据学补偿误差。
  • 实测成功率接近完美,比顶尖方法高38.4%。
  • 无需人工调教,适合真实场景快速部署。

机器人装配因需精确、接触频繁的操作而长期面临挑战。尽管基于仿真的学习已能训练出稳健的装配策略,但其在真实环境中的表现常受仿真到现实差距影响。相反,真实世界强化学习虽避免了该差距,却高度依赖人工干预且泛化能力差。本文提出SPARR:一种结合仿真训练基础策略与真实世界残差策略的混合方法。基础策略在仿真中使用低层状态观测和密集奖励训练,提供初始行为先验;残差策略在真实世界中利用视觉观测和稀疏奖励学习,弥补动力学差异与传感器噪声。大量真实实验表明,SPARR在多种两部件装配任务中实现近完美成功率。相比最先进零样本仿真到现实方法,成功率达提升38.4%,周期时间减少29.7%。此外,SPARR无需人类专家介入,优于依赖人工监督的先进真实世界强化学习方法。

原文摘要 · Abstract (English)

Robotic assembly presents a long-standing challenge due to its requirement for precise, contact-rich manipulation. While simulation-based learning has enabled the development of robust assembly policies, their performance often degrades when deployed in real-world settings due to the sim-to-real gap. Conversely, real-world reinforcement learning (RL) methods avoid the sim-to-real gap, but rely heavily on human supervision and lack generalization ability to environmental changes. In this work, we propose a hybrid approach that combines a simulation-trained base policy with a real-world residual policy to efficiently adapt to real-world variations. The base policy, trained in simulation using low-level state observations and dense rewards, provides strong priors for initial behavior. The residual policy, learned in the real world using visual observations and sparse rewards, compensates for discrepancies in dynamics and sensor noise. Extensive real-world experiments demonstrate that our method, SPARR, achieves near-perfect success rates across diverse two-part assembly tasks. Compared to the state-of-the-art zero-shot sim-to-real methods, SPARR improves success rates by 38.4% while reducing cycle time by 29.7%. Moreover, SPARR requires no human expertise, in contrast to the state-of-the-art real-world RL approaches that depend heavily on human supervision.

机器人装配强化学习仿真到现实残差策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。