arXiv:2505.05609cs.LG2025-05AAAI被引 7

提出抗数据污染的强化学习方法,提升策略稳定性。

On Corruption-Robustness in Performative Reinforcement Learning

  • 基于鲁棒均值估计与凸凹优化,处理带污染梯度的重复训练。
  • 理论证明近似稳定策略的误差与√ε线性相关。
  • 适合高噪声环境下需可靠决策的智能体应用。

在行为型强化学习中,环境奖励与转移函数依赖于智能体策略。先前研究关注重复重训练方法在有限样本下收敛至行为稳定策略。该方法反复求解凸凹目标的鞍点,以估计正则化强化学习问题的拉格朗日函数。本文旨在扩展此类方法,使其在数据受污染情况下仍能运行。具体考虑Huber的ε-污染模型,即ε比例的数据点受任意对抗噪声污染。我们提出一种基于受污染梯度的凸凹优化重复重训练方法,并引入一种针对该问题设计的鲁棒均值估计器。理论上证明,该方法实现最后一轮迭代收敛至近似稳定策略,逼近误差关于√ε呈线性关系。实验验证了在行为型强化学习中考虑数据污染的重要性。

原文摘要 · Abstract (English)

In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performative RL has studied the convergence of repeated retraining approaches to a performatively stable policy. In the finite sample regime, these approaches repeatedly solve for a saddle point of a convex-concave objective, which estimates the Lagrangian of a regularized version of the reinforcement learning problem. In this paper, we aim to extend such repeated retraining approaches, enabling them to operate under corrupted data. More specifically, we consider Huber's $ε$-contamination model, where an $ε$ fraction of data points is corrupted by arbitrary adversarial noise. We propose a repeated retraining approach based on convex-concave optimization under corrupted gradients and a novel problem-specific robust mean estimator for the gradients. We prove that our approach exhibits last-iterate convergence to an approximately stable policy, with the approximation error linear in $\sqrtε$. We experimentally demonstrate the importance of accounting for corruption in performative RL.

强化学习鲁棒性数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。