用强化学习提升垃圾箱吊装精度,解决起重机摇晃难题
Residual Reinforcement Learning for Waste-Container Lifting Using Large-Scale Cranes with Underactuated Tools
- 在已有控制基础上加学习型补偿模块,不从零训练
- 仿真中轨迹误差降低,摆动减少,成功率显著提高
- 适合需要高精度、强鲁棒性的工业自动化场景
本文研究城市环境中液压装载起重机搭载欠驱动卸料装置完成垃圾箱吊装任务的控制问题,提出一种残差强化学习(RRL)方法,将名义笛卡尔控制器与学习型残差策略相结合。所有实验在仿真中进行,任务对卸料钩与箱体环之间的几何公差要求极高,精确轨迹跟踪和抑振至关重要。名义控制器采用阻抗控制实现轨迹跟踪,结合摆动感知的阻尼策略,并通过带零空间姿态项的阻尼最小二乘逆运动学生成关节速度指令。在Isaac Lab中训练的PPO残差策略,有效补偿未建模动态和参数变化,提升了精度与鲁棒性,无需端到端从头学习。通过随机化初始状态及对载荷属性、执行器增益和被动关节参数的域随机化,增强了泛化能力。仿真结果表明,相比仅使用名义控制器,该方法在轨迹跟踪精度、振荡抑制和吊装成功率方面均有明显提升。
原文摘要 · Abstract (English)
This paper studies the container lifting phase of a waste-container recycling task in urban environments, performed by a hydraulic loader crane equipped with an underactuated discharge unit, and proposes a residual reinforcement learning (RRL) approach that combines a nominal Cartesian controller with a learned residual policy. All experiments are conducted in simulation, where the task is characterized by tight geometric tolerances between the discharge-unit hooks and the container rings relative to the overall crane scale, making precise trajectory tracking and swing suppression essential. The nominal controller uses admittance control for trajectory tracking and pendulum-aware swing damping, followed by damped least-squares inverse kinematics with a nullspace posture term to generate joint velocity commands. A PPO-trained residual policy in Isaac Lab compensates for unmodeled dynamics and parameter variations, improving precision and robustness without requiring end-to-end learning from scratch. We further employ randomized episode initialization and domain randomization over payload properties, actuator gains, and passive joint parameters to enhance generalization. Simulation results demonstrate improved tracking accuracy, reduced oscillations, and higher lifting success rates compared to the nominal controller alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。