提升机器人装配策略在复杂初始条件下的成功率。
Refinery: Active Fine-tuning and Deployment-time Optimization for Contact-Rich Policies
- 用贝叶斯优化微调单个策略,降低性能波动。
- 部署时用高斯混合模型采样初始状态,成功率提升至91.51%。
- 支持多步装配链,无需多步训练即可完成8部件组装。
基于仿真的学习已使精密、接触丰富的任务(如机器人装配)在高观测噪声和控制误差下达到约80%的成功率。尽管这一表现对研究应用可能足够,但未达工业标准,且策略串联异常脆弱。主要瓶颈在于不同初始条件下单个策略性能的高方差。本文提出Refinery框架,有效弥合性能差距,增强策略在各类初始条件下的鲁棒性。方法包括:利用贝叶斯优化引导的微调以改进个体策略;部署时采用高斯混合模型采样初始状态,选择最有利于执行成功的起始点。实验表明,与现有最优方法相比,该框架在仿真中平均成功率提升10.98%,达到91.51%,且真实世界表现相当。此外,经微调的策略可成功串联,实现长时程、多部件装配,无需显式多步训练即完成最多8部件的组装。
原文摘要 · Abstract (English)
Simulation-based learning has enabled policies for precise, contact-rich tasks (e.g., robotic assembly) to reach high success rates (~80%) under high levels of observation noise and control error. Although such performance may be sufficient for research applications, it falls short of industry standards and makes policy chaining exceptionally brittle. A key limitation is the high variance in individual policy performance across diverse initial conditions. We introduce Refinery, an effective framework that bridges this performance gap, robustifying policy performance across initial conditions. We propose Bayesian Optimization-guided fine-tuning to improve individual policies, and Gaussian Mixture Model-based sampling during deployment to select initializations that maximize execution success. Using Refinery, we improve mean success rates by 10.98% over state-of-the-art methods in simulation-based learning for robotic assembly, reaching 91.51% in simulation and comparable performance in the real world. Furthermore, we demonstrate that these fine-tuned policies can be chained to accomplish long-horizon, multi-part assembly$\unicode{x2013}$successfully assembling up to 8 parts without requiring explicit multi-step training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。