在高维稀疏强化学习中,提出抗数据污染的鲁棒算法,首次实现非平凡的理论保证。
Sparse Offline Reinforcement Learning with Corruption Robustness
- 结合稀疏性设计新算法,避免传统悲观奖励的过度保守问题。
- 在单策略集中覆盖下,实现首个非平凡的理论保证,样本数小于特征维数仍有效。
- 适用于高维稀疏马尔可夫决策过程,对强数据污染具有鲁棒性,适合实际部署场景。
我们研究了在离线稀疏强化学习中对强数据污染的鲁棒性。在该设置下,对手可任意扰动部分从高维但稀疏的马尔可夫决策过程收集的轨迹,目标是估计一个近似最优策略。主要挑战在于:当样本数 $N$ 小于特征维度 $d$ 时,必须利用稀疏性才能获得非平凡的保证,但这一方向在离线强化学习中尚未系统研究。我们在均匀覆盖和稀疏单策略集中性假设下分析该问题。虽然标准的最小二乘值迭代(LSVI)在均匀覆盖下表现良好,但我们发现将稀疏性融入LSVI不自然,其分析可能因过于悲观的奖励而失效。为此,我们提出了基于稀疏鲁棒估计器预言的演员-评论家方法,避免使用逐点悲观奖励,在单策略集中性覆盖下首次提供稀疏离线强化学习的非平凡保证。此外,我们将结果扩展到污染设置,证明该算法在强污染下依然鲁棒。我们的结果首次在高维稀疏MDP与单策略集中性覆盖及污染条件下提供非平凡保证,表明在传统鲁棒离线强化学习技术可能失效的区域,仍可学习近似最优策略。
原文摘要 · Abstract (English)
We investigate robustness to strong data corruption in offline sparse reinforcement learning (RL). In our setting, an adversary may arbitrarily perturb a fraction of the collected trajectories from a high-dimensional but sparse Markov decision process, and our goal is to estimate a near optimal policy. The main challenge is that, in the high-dimensional regime where the number of samples $N$ is smaller than the feature dimension $d$, exploiting sparsity is essential for obtaining non-vacuous guarantees but has not been systematically studied in offline RL. We analyse the problem under uniform coverage and sparse single-concentrability assumptions. While Least Square Value Iteration (LSVI), a standard approach for robust offline RL, performs well under uniform coverage, we show that integrating sparsity into LSVI is unnatural, and its analysis may break down due to overly pessimistic bonuses. To overcome this, we propose actor-critic methods with sparse robust estimator oracles, which avoid the use of pointwise pessimistic bonuses and provide the first non-vacuous guarantees for sparse offline RL under single-policy concentrability coverage. Moreover, we extend our results to the contaminated setting and show that our algorithm remains robust under strong contamination. Our results provide the first non-vacuous guarantees in high-dimensional sparse MDPs with single-policy concentrability coverage and corruption, showing that learning a near-optimal policy remains possible in regimes where traditional robust offline RL techniques may fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。