用强化学习优化差速机器人控制器参数,提升效率与稳定性。
Gray-Box Computed Torque Control for Differential-Drive Mobile Robot Tracking
- 用灰盒控制器替代黑箱策略网络,结合物理先验知识。
- 仅需少量训练样本即可找到最优控制参数,响应时间临界阻尼。
- 适合需要高效稳定控制的移动机器人研发人员。
本研究提出一种基于学习的非线性控制算法,用于差速移动机器人的轨迹跟踪。传统计算力矩法(CTM)因系统参数不准确而效果受限,深度强化学习(DRL)则存在样本效率低、闭环稳定性差的问题。本文提出的灰盒计算力矩控制器(CTC)将DRL代理的黑箱策略网络替换为具有物理可解释性的控制器,显著提升样本效率并保证闭环稳定性。通过使用Twin-Delayed Deep Deterministic Policy Gradient(TD3)算法,仅需少数短时学习回合即可为任意奖励函数寻得最优控制器参数。部分控制器参数被限制在已知合理范围内,确保学习值符合物理实际;同时引入技术使闭环响应达到临界阻尼。控制器在MuJoCo物理引擎中模拟的差速机器人上进行评估,性能优于原始CTC和传统运动学控制器。
原文摘要 · Abstract (English)
This study presents a learning-based nonlinear algorithm for tracking control of differential-drive mobile robots. The Computed Torque Method (CTM) suffers from inaccurate knowledge of system parameters, while Deep Reinforcement Learning (DRL) algorithms are known for sample inefficiency and weak stability guarantees. The proposed method replaces the black-box policy network of a DRL agent with a gray-box Computed Torque Controller (CTC) to improve sample efficiency and ensure closed-loop stability. This approach enables finding an optimal set of controller parameters for an arbitrary reward function using only a few short learning episodes. The Twin-Delayed Deep Deterministic Policy Gradient (TD3) algorithm is used for this purpose. Additionally, some controller parameters are constrained to lie within known value ranges, ensuring the RL agent learns physically plausible values. A technique is also applied to enforce a critically damped closed-loop time response. The controller's performance is evaluated on a differential-drive mobile robot simulated in the MuJoCo physics engine and compared against the raw CTC and a conventional kinematic controller.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。