用物理感知网络加速多旋翼控制策略训练,节省超七成环境交互。
Efficient Knowledge Transfer for Jump-Starting Control Policy Learning of Multirotors through Physics-Aware Neural Architectures
- 构建物理感知神经架构,融合强化学习与监督分配网络。
- 跨构型初始化可减少73.5%环境交互,提升训练效率。
- 适合需快速部署多旋翼控制的机器人研究者与工程师。
高效训练机器人控制策略是重大挑战,通过跨机体知识迁移可显著受益。本文提出基于库的初始化方案,实现多旋翼配置间的有效知识复用。利用融合强化学习控制器与监督控制分配网络的物理感知神经架构,支持已训练策略的重用。通过基于策略评估的相似性度量,从库中识别适配的初始策略。实验表明该度量与达到目标性能所需环境交互次数的减少高度相关,具备良好适用性。仿真与真实世界实验均验证:该架构达当前最优控制性能,且在多种四旋翼与六旋翼设计上,平均节省73.5%环境交互(相较从零训练),为强化学习中的跨机体迁移提供了高效路径。
原文摘要 · Abstract (English)
Efficiently training control policies for robots is a major challenge that can greatly benefit from utilizing knowledge gained from training similar systems through cross-embodiment knowledge transfer. In this work, we focus on accelerating policy training using a library-based initialization scheme that enables effective knowledge transfer across multirotor configurations. By leveraging a physics-aware neural control architecture that combines a reinforcement learning-based controller and a supervised control allocation network, we enable the reuse of previously trained policies. To this end, we utilize a policy evaluation-based similarity measure that identifies suitable policies for initialization from a library. We demonstrate that this measure correlates with the reduction in environment interactions needed to reach target performance and is therefore suited for initialization. Extensive simulation and real-world experiments confirm that our control architecture achieves state-of-the-art control performance, and that our initialization scheme saves on average up to $73.5\%$ of environment interactions (compared to training a policy from scratch) across diverse quadrotor and hexarotor designs, paving the way for efficient cross-embodiment transfer in reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。