用多样化重置让机器人学会复杂操作,无需人工设计奖励或教程
Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
- 通过程序化重置让强化学习接触更多机器人-物体交互场景
- 在长时序操作任务上成功率显著超越现有方法,且性能随算力持续提升
- 可零样本迁移到真实机器人,具备自主重试能力,适合通用灵巧操作研究
大规模并行物理仿真中的强化学习推动了仿真到现实的机器人学习进展。然而,当前方法仍脆弱且任务特定,依赖大量针对每项任务的人工工程来设计奖励函数、训练课程和示范。即使如此,它们在长时序、高接触性操作任务中仍经常失败,且性能随算力增长迅速饱和,因训练反复集中在状态空间的狭窄区域。我们提出OmniReset,一种简单且可扩展的框架,使基于策略的强化学习能仅用单一奖励函数、固定超参数、无训练课程和无人类示范,稳健解决一系列灵巧操作任务。核心洞察是:通过模拟器重置系统性暴露强化学习算法于灵巧操作所依赖的多样机器人-物体交互。OmniReset以极低人工干预生成此类重置,将额外算力直接转化为更广的行为覆盖和持续性能提升。我们验证了OmniReset能优雅扩展至现有方法无法应对的长时序灵巧操作任务,并在远更宽的初始条件范围内学习出鲁棒策略。最终,我们将OmniReset提炼为视觉-运动策略,在零样本迁移至真实世界时表现出更强的重试能力和显著更高的成功率。
原文摘要 · Abstract (English)
Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engineering to design rewards, curricula, and demonstrations. Even with this engineering, they often fail on long-horizon, contact-rich manipulation tasks and do not meaningfully scale with compute, as performance quickly saturates when training revisits the same narrow regions of state space. We introduce OmniReset, a simple and scalable framework that enables on-policy reinforcement learning to robustly solve a broad class of dexterous manipulation tasks using a single reward function, fixed algorithm hyperparameters, no curricula, and no human demonstrations. Our key insight is that long-horizon exploration can be dramatically simplified by using simulator resets to systematically expose the RL algorithm to the diverse set of robot-object interactions which underlie dexterous manipulation. OmniReset programmatically generates such resets with minimal human input, converting additional compute directly into broader behavioral coverage and continued performance gains. We show that OmniReset gracefully scales to long-horizon dexterous manipulation tasks beyond the capabilities of existing approaches and is able to learn robust policies over significantly wider ranges of initial conditions than baselines. Finally, we distill OmniReset into visuomotor policies which display robust retrying behavior and substantially higher success rates than baselines when transferred to the real world zero-shot. Project webpage: https://weirdlabuw.github.io/omnireset/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。