Transformer可实现非平稳强化学习中的近最优动态遗憾。
Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
- 利用上下文学习逼近非平稳环境策略
- 理论证明可达到近最优动态遗憾边界
- 实验表现优于或媲美现有专家算法
Transformer在多个领域展现出卓越性能。尽管其在上下文学习中进行强化学习的能力已从理论和实证层面得到验证,但在非平稳环境中的行为仍不明确。本文填补了这一空白,证明Transformer可在非平稳设置下实现近乎最优的动态遗憾。我们证明Transformer能够近似用于处理非平稳环境的策略,并可在上下文学习框架中学习该近似器。实验结果进一步表明,Transformer在这些环境中可匹配甚至超越现有专家算法的表现。
原文摘要 · Abstract (English)
Transformers have demonstrated exceptional performance across a wide range of domains. While their ability to perform reinforcement learning in-context has been established both theoretically and empirically, their behavior in non-stationary environments remains less understood. In this study, we address this gap by showing that transformers can achieve nearly optimal dynamic regret bounds in non-stationary settings. We prove that transformers are capable of approximating strategies used to handle non-stationary environments and can learn the approximator in the in-context learning setup. Our experiments further show that transformers can match or even outperform existing expert algorithms in such environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。