用强化学习让挖机自动抓石头,不依赖物理建模也能稳定完成。
Learning to Capture Rocks using an Excavator: A Reinforcement Learning Approach with Guiding Reward Formulation
- 基于PPO算法的无模型强化学习,直接输出挖机关节速度指令。
- 在随机化环境下训练,对未知石头和土壤条件成功率高,接近真人水平。
- 无需特殊夹具或材料模型,适合真实工地应用。
使用标准挖斗抓取岩石是一项极具挑战性的任务,通常需要经验丰富的操作员。与挖掘土壤不同,该任务涉及在非结构化环境中操纵大而形状不规则的岩石,复杂的颗粒材料接触交互使得基于模型的控制方法难以实现。现有自主挖掘方法主要针对连续介质或依赖专用夹具,限制了其在真实施工场地的应用。本文提出一种完全数据驱动的控制框架,无需显式建模岩石或土壤属性。在AGX Dynamics仿真器中,采用近端策略优化(PPO)算法训练一个无模型强化学习代理,并设计引导式奖励函数。学习到的策略直接输出卡特365挖机模型的臂、铲杆和铲斗的关节速度指令。通过大量随机化岩石几何、密度、质量以及挖斗、岩石和目标位置的初始配置,提升了鲁棒性。据我们所知,这是首个针对岩石抓取任务开发并评估的基于强化学习的控制器。实验结果表明,该策略能良好泛化至未见岩石和变化的土壤条件,在保持机器稳定性的同时,成功率与人类参与者相当。这些发现证明了无需专用硬件或详细材料模型,即可实现基于学习的离散物体挖掘策略的可行性。
原文摘要 · Abstract (English)
Rock capturing with standard excavator buckets is a challenging task typically requiring the expertise of skilled operators. Unlike soil digging, it involves manipulating large, irregular rocks in unstructured environments where complex contact interactions with granular material make model-based control impractical. Existing autonomous excavation methods focus mainly on continuous media or rely on specialized grippers, limiting their applicability to real-world construction sites. This paper introduces a fully data-driven control framework for rock capturing that eliminates the need for explicit modeling of rock or soil properties. A model-free reinforcement learning agent is trained in the AGX Dynamics simulator using the Proximal Policy Optimization (PPO) algorithm and a guiding reward formulation. The learned policy outputs joint velocity commands directly to the boom, arm, and bucket of a CAT365 excavator model. Robustness is enhanced through extensive domain randomization of rock geometry, density, and mass, as well as the initial configurations of the bucket, rock, and goal position. To the best of our knowledge, this is the first study to develop and evaluate an RL-based controller for the rock capturing task. Experimental results show that the policy generalizes well to unseen rocks and varying soil conditions, achieving high success rates comparable to those of human participants while maintaining machine stability. These findings demonstrate the feasibility of learning-based excavation strategies for discrete object manipulation without requiring specialized hardware or detailed material models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。