用动态实例集提升机器人抓取的零样本迁移能力
Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching

- 引入动态实例集同时模拟多种物理不确定性
- 仅10个实例即实现无需实机微调的可靠抓取
- 适合高精度灵巧操作与仿真到现实迁移研究
灵巧操作依赖物理规律且对建模误差和感知噪声极为敏感,导致仿真到现实的迁移极具挑战。传统领域随机化(DR)每轮只随机一个实例,难以充分覆盖真实动态的变异性。为此,我们提出域随机实例集(DRIS),同时表征并传播一组随机实例,更丰富地逼近不确定动态,使策略学会应对多种可能结果。理论分析表明,DRIS可生成更鲁棒的策略,显著减少对真实世界微调的需求,即使仅使用少量实例(如10个)亦可。我们在一个具有挑战性的实时反应抓取任务上验证了该方法。不同于使用机械稳定结构(如曲面或封闭结构)的传统抓取器,我们的系统采用无被动稳定功能的平板,任务对噪声极度敏感,需快速反应动作。学习到的策略展现出强鲁棒性,在零样本条件下成功实现仿真到现实的迁移。
原文摘要 · Abstract (English)
Dexterous manipulation is physics-intensive and highly sensitive to modeling errors and perception noise, making sim-to-real transfer prohibitively challenging. Domain randomization (DR) is commonly used to improve the robustness of learned policies for such tasks, but conventional DR randomizes one instance per episode, offering very limited exposure to the variability of real-world dynamics. To this end, we propose Domain-Randomized Instance Set (DRIS), which represents and propagates a set of randomized instances simultaneously, providing richer approximation of uncertain dynamics and enabling policies to learn actions that account for multiple possible outcomes. Supported by theoretical analysis, we show that DRIS yields more robust policies and alleviates the need for real-world fine-tuning, even with a modest number of instances (e.g., 10). We demonstrate this on a challenging reactive catching task. Unlike traditional catching setups that use end-effectors designed to mechanically stabilize the object (e.g., curved or enclosing surfaces), our system uses a flat plate that offers no passive stabilization, making the task highly sensitive to noise and requiring rapid reactive motions. The learned policies exhibit strong robustness to uncertainties and achieve reliable zero-shot sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。