提出抗干扰强化学习算法,实现异步数据下近最优的鲁棒性。
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
- 设计新Q-learning变体,对抗恶意污染奖励信号
- 在异步相关数据中实现接近最优的有限时间收敛率
- 无需了解奖励分布,首次给出异步Q-learning的鲁棒性保障
我们研究在折扣无限时域强化学习设置下,存在对抗性污染奖励时学习最优策略的问题。为此,我们提出一种新颖的鲁棒Q-learning算法,并在具有时间相关数据的挑战性异步采样模型下进行分析。尽管存在污染,我们的方法在有限时间内保证的性能与现有结果相当,仅多出一个与污染样本比例成正比的附加项。我们还建立了信息论下界,表明我们的保证是近最优的。值得注意的是,该算法对底层奖励分布完全无感,是首个为异步Q-learning提供有限时间鲁棒性保证的方法。分析中的关键创新是一种针对几乎鞅的改进版Azuma-Hoeffding不等式,可能在强化学习算法研究中具有更广泛应用。
原文摘要 · Abstract (English)
We study the problem of learning the optimal policy in a discounted, infinite-horizon reinforcement learning (RL) setting in the presence of adversarially corrupted rewards. To address this problem, we develop a novel robust variant of the \(Q\)-learning algorithm and analyze it under the challenging asynchronous sampling model with time-correlated data. Despite corruption, we prove that the finite-time guarantees of our approach match existing bounds, up to an additive term that scales with the fraction of corrupted samples. We also establish an information-theoretic lower bound, revealing that our guarantees are near-optimal. Notably, our algorithm is agnostic to the underlying reward distribution and provides the first finite-time robustness guarantees for asynchronous \(Q\)-learning. A key element of our analysis is a refined Azuma-Hoeffding inequality for almost-martingales, which may have broader applicability in the study of RL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。