用数据自适应优化提升自行车自动驾驶的平衡能力
An Adaptive Data-Enabled Policy Optimization Approach for Autonomous Bicycle Control
- 内外环协同:内环用反馈线性化稳定系统,外环用数据驱动策略优化
- 实测表明:追踪参考倾角和倾角速率更精准,优于纯反馈线性化方法
- 适合研究机器人控制、自适应学习的工程师与学者
本文提出一种统一控制框架,将反馈线性化(FL)控制器置于内环以稳定并部分线性化本就非线性且不稳定的自主自行车系统,同时在外部环路引入自适应数据使能策略优化(DeePO)控制器以增强适应性与鲁棒性。初始控制策略由一组离线的、持续激励的输入与状态数据获得。为提升稳定性并补偿系统非线性及扰动,采用促进鲁棒性的正则化项优化初始策略;同时通过遗忘因子改进DeePO的自适应能力,以应对时变动态特性。所提出的DeePO+FL方法在仿真与真实仪器化自主自行车实验中得到验证,结果表明其在参考倾角与倾角速率跟踪精度上显著优于仅使用FL的方法。
原文摘要 · Abstract (English)
This paper presents a unified control framework that integrates a Feedback Linearization (FL) controller in the inner loop with an adaptive Data-Enabled Policy Optimization (DeePO) controller in the outer loop to balance an autonomous bicycle. While the FL controller stabilizes and partially linearizes the inherently unstable and nonlinear system, its performance is compromised by unmodeled dynamics and time-varying characteristics. To overcome these limitations, the DeePO controller is introduced to enhance adaptability and robustness. The initial control policy of DeePO is obtained from a finite set of offline, persistently exciting input and state data. To improve stability and compensate for system nonlinearities and disturbances, a robustness-promoting regularizer refines the initial policy, while the adaptive section of the DeePO framework is enhanced with a forgetting factor to improve adaptation to time-varying dynamics. The proposed DeePO+FL approach is evaluated through simulations and real-world experiments on an instrumented autonomous bicycle. Results demonstrate its superiority over the FL-only approach, achieving more precise tracking of the reference lean angle and lean rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。