用可微分动力学实现机器人控制器与扰动观测器联合自动调参。
MetaTune: Adjoint-based Meta-tuning via Robotic Differentiable Dynamics
- 基于可微分系统动力学,通过伴随法反向计算元梯度。
- 在四旋翼任务中追踪误差降低15%~40%,梯度计算提速超50%。
- 适合需要高鲁棒性控制的复杂机器人系统部署。
基于扰动观测器的控制在增强机器人系统对不确定性的鲁棒性方面展现出潜力。然而,由于控制器增益与观测器参数之间存在强耦合,调参仍具挑战。本文提出MetaTune,一种通过可微分闭环元学习联合自动调参反馈控制器与扰动观测器的统一框架。MetaTune融合可移植神经策略与源自可微分系统动力学的物理信息梯度,实现在不同任务与工况下的自适应增益调整。我们开发了一种伴随方法,可高效地沿时间反向计算关于自适应增益的元梯度,直接最小化代价到目标。相比现有前向方法,本方法将计算复杂度降至数据时长的线性。在四旋翼控制任务中,MetaTune实现竞争力或更优的追踪性能,同时将梯度计算时间减少超过50%。在PX4-Gazebo硬件在环仿真中,学习到的策略实现零样本迁移,在激进飞行下追踪均方根误差(RMSE)降低约15–20%,强干扰下最高降低40%。
原文摘要 · Abstract (English)
Disturbance observer-based control has shown promise in robustifying robotic systems against uncertainties. However, tuning such systems remains challenging due to the strong coupling between controller gains and observer parameters. In this work, we propose MetaTune, a unified framework for joint auto-tuning of feedback controllers and disturbance observers through differentiable closed-loop meta-learning. MetaTune integrates a portable neural policy with physics-informed gradients derived from differentiable system dynamics, enabling adaptive gains across tasks and operating conditions. We develop an adjoint method that efficiently computes the meta-gradients with respect to adaptive gains backward in time to directly minimize the cost-to-go. Compared to existing forward methods, our approach reduces the computational complexity to be linear in the data horizon. On quadrotor control tasks, MetaTune achieves competitive or improved tracking performance while reducing gradient computation time by more than 50\%. In PX4-Gazebo hardware-in-the-loop simulation, the learned policy transfers zero-shot and reduces tracking RMSE by about 15--20\% in aggressive flight and up to 40\% under strong disturbances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。