用粒子模拟+分阶段训练,让挖掘机自动挖掉埋得越来越深的障碍物。
Autonomous Obstacle Removal for Excavators through Policy Learning with Particle Simulation

- 根据障碍物埋深动态调整模拟难度和粒子数量,降低训练成本。
- 3天内学会挖障碍,真实12吨挖掘机也能成功应用。
- 适合做智能工程机械自主作业研究的人看。
自主清除地面障碍是土方作业的重要任务,但因土壤-障碍物状态随重复挖掘不断变化,难以自动化。学习这种状态依赖行为需要能再现累积相互作用的训练环境,包括接触状态、地形变形和障碍可见性。粒子模拟适合此类策略学习,但计算开销大,重复挖掘周期进一步增加学习成本。我们发现障碍物埋深同时决定任务难度与模拟成本:埋得越深,移除越难,所需粒子越多。据此提出一种基于埋深的课程学习策略。构建一个高效仿真到现实的策略学习框架,策略通过RGB-D感知地形与障碍信息,输出参数化挖掘轨迹;模拟器在真实挖掘机上复现相同的观测-动作接口,在可控埋深条件下运行。课程从浅埋开始,逐步加深埋深并调整粒子数,同步控制任务难度与模拟成本。实验表明,该框架成功学习出有效移除策略,而基线方法即使训练一周仍失败。所提课程在三天内即达良好性能,并成功迁移至真实12吨挖掘机在开放地面对多种钢制障碍物的作业,证明了其鲁棒性。
原文摘要 · Abstract (English)
Autonomous obstacle removal from the ground is an important earthwork task, but this is difficult to automate because an excavator must adapt its excavation trajectories over repeated cycles as soil-obstacle conditions change. Learning such state-dependent behavior requires a training environment that reproduces accumulated soil-obstacle interactions, including contact states, terrain deformation, and obstacle visibility. Accordingly, particle-based simulation is suitable for the relevant policy learning. However, particle simulation is computationally expensive, and repeated excavation cycles further increase the learning cost. We observe that the burial condition of an obstacle governs both task difficulty and simulation cost: deeper burial makes obstacle removal harder while also requiring more particles for accurate simulation. This observation motivates a burial-conditioned curriculum learning strategy. We propose a time-efficient sim-to-real policy learning framework in which the policy observes terrain and obstacle information from RGB-D measurements and then outputs a parameterized excavation trajectory; in this process, the simulator reproduces in a real-world excavator the same observation-action interface it uses under controllable burial conditions. The curriculum begins with shallow burial conditions and progressively increases burial depth while adjusting particle count, thus simultaneously controlling task difficulty and simulation cost. Experiments show that the proposed framework successfully learns an effective obstacle-removal policy, whereas baseline methods fail even after a full week of training. The proposed curriculum achieves effective performance within three days and achieves successful transfer to a real 12-ton excavator operating on open ground with various steel obstacles, thus demonstrating robust obstacle removal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。