arXiv:2411.14117cs.LGcs.AI2024-11被引 5

用物理启发方法提升强化学习在复杂问题中的效率与探索能力

Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems

  • 结合蒙特卡洛采样与最优控制,通过神经网络实现高效策略优化
  • 在稀疏奖励、状态陷阱等难题上,计算效率优于现有最先进算法
  • 多智能体协同+熵奖励机制,实现自适应探索与利用平衡

我们提出一种新颖的、计算高效的强化学习方法,用于解决困难的非线性问题。该方法将计算物理/化学中的伞形采样与最优控制相结合,并基于神经网络和策略梯度实现。在处理具有稀疏奖励、状态陷阱及无终止状态的困难强化学习问题时,其计算效率和实现通用性均优于所有现有最先进算法。该方法采用一组同时行动的智能体,通过引入包含群体熵的修正奖励函数,实现最优的探索-利用平衡。

原文摘要 · Abstract (English)

We report a novel, computationally efficient approach for solving hard nonlinear problems of reinforcement learning (RL). Here we combine umbrella sampling, from computational physics/chemistry, with optimal control methods. The approach is realized on the basis of neural networks, with the use of policy gradient. It outperforms, by computational efficiency and implementation universality, all available state-of-the-art algorithms, in application to hard RL problems with sparse reward, state traps and lack of terminal states. The proposed approach uses an ensemble of simultaneously acting agents, with a modified reward which includes the ensemble entropy, yielding an optimal exploration-exploitation balance.

强化学习高效算法多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。