用最小注意力机制提升强化学习的快速适应与稳定性。
Meta-reinforcement learning with minimum attention
- 将最小作用原理引入强化学习奖励设计,实现高效控制。
- 在高维非线性系统中实现少样本快速适应与抗扰动能力。
- 兼顾能量效率,适合对稳定性与能效要求高的场景。
最小注意力将最小作用原理应用于状态与时间变化的控制,首次由Brockett提出,其正则化项在模拟生物控制(如运动学习)中具有重要意义。本文将其作为奖励的一部分引入强化学习,探究其与元学习和稳定性的关联。针对高维非线性动力学,采用基于集成的模型学习与基于梯度的元策略学习交替进行。实验表明,最小注意力在少样本快速适应与模型及环境扰动下的方差抑制方面优于当前最先进模型无关与模型相关的强化学习算法,同时展现出更高的能量效率。
原文摘要 · Abstract (English)
Minimum attention applies the least action principle to changes of control concerning state and time, first proposed by Brockett. The involved regularization is highly relevant in emulating biological control, such as motor learning. We apply minimum attention in reinforcement learning (RL) as part of the rewards and investigate its connection to meta-learning and stabilization. Specifically, model-based meta-learning with minimum attention is explored in high-dimensional nonlinear dynamics. Ensemble-based model learning and gradient-based meta-policy learning are alternately performed. Empirically, the minimum attention does show outperforming competence in comparison to the state-of-the-art algorithms of model-free and model-based RL, i.e., fast adaptation in few shots and variance reduction from the perturbations of the model and environment. Furthermore, the minimum attention demonstrates an improvement in energy efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。