提出一种新强化学习方法,用理论保证提升探索效率。
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
- 从对偶优化视角设计新算法,统一优化探索与利用
- 在线性马尔可夫决策过程下实现近最优后悔率
- 适合需要高效探索的复杂模型在线学习场景
使用复杂函数逼近(如Transformer和深度神经网络)的在线强化学习在现代人工智能中扮演重要角色。尽管应用广泛,但如何在探索与利用之间取得平衡仍是长期挑战,尤其缺乏兼具高效性与理论保证的实际方案。受乐观正则化启发,本文从原始-对偶优化角度重新诠释了乐观原则。基于此,提出一种新的价值激励演员-评论家(VAC)方法,通过单一易优化目标同时优化探索与利用:所学的状态-动作值和策略估计既符合采集的数据转移,又能带来更高的价值函数。理论上,该方法在有限与无限时域的线性马尔可夫决策过程(MDPs)下具有近最优后悔率,且在合理假设下可推广至一般函数逼近设定。
原文摘要 · Abstract (English)
Online reinforcement learning (RL) with complex function approximations such as transformers and deep neural networks plays a significant role in the modern practice of artificial intelligence. Despite its popularity and importance, balancing the fundamental trade-off between exploration and exploitation remains a long-standing challenge; in particular, we are still in lack of efficient and practical schemes that are backed by theoretical performance guarantees. Motivated by recent developments in exploration via optimistic regularization, this paper provides an interpretation of the principle of optimism through the lens of primal-dual optimization. From this fresh perspective, we set forth a new value-incentivized actor-critic (VAC) method, which optimizes a single easy-to-optimize objective integrating exploration and exploitation -- it promotes state-action and policy estimates that are both consistent with collected data transitions and result in higher value functions. Theoretically, the proposed VAC method has near-optimal regret guarantees under linear Markov decision processes (MDPs) in both finite-horizon and infinite-horizon settings, which can be extended to the general function approximation setting under appropriate assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。