用视觉+分层学习让重达6吨的机器人精准自主导航
Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework
- 分层框架融合视觉定位、强化学习规划与自适应控制
- 6000公斤机器人在复杂地形上定位误差仅3-4厘米
- 故障后能自动安全返航,适合重型移动机器人应用
强化学习在机器人领域潜力巨大,但基于探索的训练难以保障大型机器人安全部署。本文提出一种新型分层目标到达框架,集成立体视觉位姿估计、约束强化学习运动规划、执行器级鲁棒自适应控制(RAC)和监督式安全返回逻辑。立体视觉实现实时位姿估计,支持回环闭合、地图融合与重定位。强化学习规划器通过问题特异性奖励结构与运动约束,生成平滑可行的目标路径,促进目标进展、减少振荡、保持视觉一致性并符合重型滑移转向机器人的机械限制。执行层采用缩放共轭梯度(SCG)训练的深度神经网络(DNN),近似从轮速数据到期望控制输入的准静态前馈映射。该映射结合基于对数屏障的RAC,补偿残余建模误差、打滑扰动及名义映射与实际执行器响应间的有界偏差。在轮轨跟踪子系统中,建立了有界不确定性下一致最终有界且指数收敛至依赖扰动的残留集的跟踪性能。对数安全监控器实时监测运行状态,检测故障与定位不一致,并触发安全返航模式。在6000公斤机器人于沥青与松软土壤地形上的实验表明,最终位置均方根误差约为3–4厘米,能准确跟踪强化学习生成的指令,执行器级性能优于两项基准方法,且在注入故障后成功实现自主恢复。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has strong potential in robotics, but exploration-based training complicates safe deployment on large-scale robots. For such applications, this paper proposes a novel hierarchical goal-reaching framework that integrates stereo visual pose estimation, constrained RL-based motion planning, actuator-level robust adaptive control (RAC), and supervisory safe-return logic. Stereo visual localization is used as the real-time pose-estimation interface with loop closing, map fusion, and relocalization. The RL planner generates smooth, feasible goal-reaching references using a problem-specific reward structure and motion constraints that promote goal progress, reduce oscillations, preserve vision-consistent smoothness, and respect the mechanical limits of a heavy skid-steered robot. At the actuation layer, a scaled conjugate-gradient (SCG)-trained deep neural network (DNN) approximates a quasi-static actuator feedforward map from wheel-speed data to nominal control input. This feedforward map is combined with a logarithmic-barrier-based RAC to compensate for residual modeling errors, slip-induced disturbances, and bounded mismatch between the nominal map and real actuator response. For the actuator-level wheel-tracking subsystem, uniformly ultimately bounded tracking with exponential convergence to a disturbance-dependent residual set is established under bounded uncertainty. A logarithmic safety supervisor monitors execution, detects unsafe operating conditions, including faults and localization inconsistencies, and switches the robot to safe-return mode. Experiments on a 6000 kg robot over asphalt and loose-soil terrain demonstrate approximately 3--4 cm final-position root mean square error (RMSE), accurate tracking of RL-generated commands, improved actuator-level performance over two RAC baselines, and successful autonomous recovery after fault injection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。