arXiv:2412.07477cs.ROcs.LG2024-12中稿 · IEEE Transactions …被引 4

用粗分辨率模拟预训练,加速高精度挖掘机自主学习

Progressive-Resolution Policy Distillation: Leveraging Coarse-Resolution Simulations for Time-Efficient Fine-Resolution Policy Learning

  • 分阶段迁移策略:从粗到细逐步传递政策,避免领域差距
  • 采样时间降至1/7,任务成功率与全精细模拟相当
  • 适合需要快速训练的机器人控制场景,如工程作业自动化

在土方和建筑施工中,挖掘机常需应对混杂大石块与多种土壤条件的情况,依赖熟练操作员。本文提出一种基于强化学习(RL)的自主挖掘框架,利用岩石挖掘仿真器实现。仿真中分辨率由土体空间内粒子大小/数量定义:精细分辨率能接近真实行为但计算耗时长、样本收集困难;粗分辨率虽快但偏离真实。为结合两者优势,我们探索在粗分辨率仿真中训练的策略用于精细仿真中的预训练。为此,提出新型策略学习框架——渐进式分辨率策略蒸馏(PRPD),通过中间分辨率仿真逐步转移策略,并采用保守策略迁移以避免因领域差异导致迁移失败。在岩石挖掘仿真器及九个真实岩石环境中的验证表明,PRPD将采样时间减少至不足1/7,同时保持与精细分辨率直接学习相当的任务成功率。更多视频与补充结果见项目页:https://yuki-kadokawa.github.io/prpd/

原文摘要 · Abstract (English)

In earthwork and construction, excavators often encounter large rocks mixed with various soil conditions, requiring skilled operators. This paper presents a framework for achieving autonomous excavation using reinforcement learning (RL) through a rock excavation simulator. In the simulation, resolution can be defined by the particle size/number in the whole soil space. Fine-resolution simulations closely mimic real-world behavior but demand significant calculation time and challenging sample collection, while coarse-resolution simulations enable faster sample collection but deviate from real-world behavior. To combine the advantages of both resolutions, we explore using policies developed in coarse-resolution simulations for pre-training in fine-resolution simulations. To this end, we propose a novel policy learning framework called Progressive-Resolution Policy Distillation (PRPD), which progressively transfers policies through some middle-resolution simulations with conservative policy transfer to avoid domain gaps that could lead to policy transfer failure. Validation in a rock excavation simulator and nine real-world rock environments demonstrated that PRPD reduced sampling time to less than 1/7 while maintaining task success rates comparable to those achieved through policy learning in a fine-resolution simulation. Additional videos and supplementary results are available on our project page: https://yuki-kadokawa.github.io/prpd/

强化学习机器人控制模拟训练策略蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。