在线优化无人机控制器参数,快速适应动态变化环境。
Fast Non-Episodic Adaptive Tuning of Robot Controllers with Online Policy Optimization
- 基于单轨迹模型的在线策略优化算法M-GAPS,改进状态与策略空间表示。
- 硬件实验表明,M-GAPS比基线更快找到近优参数,尤其在长周期场景下。
- 可快速应对风扰与负载变化,适用于多种机器人系统。
我们研究在线算法以调整机器人控制器参数,针对动力学、策略类别和最优目标均随时间变化的场景。系统沿单一轨迹运行,无分段或状态重置,且时变信息事先未知。以非线性几何四旋翼控制器为测试案例,提出一种实用的单轨迹模型驱动在线策略优化算法M-GAPS,结合四旋翼状态空间与策略类别的重参数化,改善优化景观。在硬件实验中,与引入人工分段的模型基于和模型无关基线对比,M-GAPS在非理想分段长度条件下更快速找到近优参数。同时,M-GAPS能迅速适应强未建模风扰与负载扰动,并在1:6缩比阿克曼转向车辆上实现类似显著性能提升。结果表明,此类在线策略优化算法具备硬件可行性,相比经典自适应控制更具灵活性,又比模型无关强化学习更稳定、数据高效。
原文摘要 · Abstract (English)
We study online algorithms to tune the parameters of a robot controller in a setting where the dynamics, policy class, and optimality objective are all time-varying. The system follows a single trajectory without episodes or state resets, and the time-varying information is not known in advance. Focusing on nonlinear geometric quadrotor controllers as a test case, we propose a practical implementation of a single-trajectory model-based online policy optimization algorithm, M-GAPS,along with reparameterizations of the quadrotor state space and policy class to improve the optimization landscape. In hardware experiments,we compare to model-based and model-free baselines that impose artificial episodes. We show that M-GAPS finds near-optimal parameters more quickly, especially when the episode length is not favorable. We also show that M-GAPS rapidly adapts to heavy unmodeled wind and payload disturbances, and achieves similar strong improvement on a 1:6-scale Ackermann-steered car. Our results demonstrate the hardware practicality of this emerging class of online policy optimization that offers significantly more flexibility than classic adaptive control, while being more stable and data-efficient than model-free reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。