arXiv:2506.06891cs.LGcs.CR2025-06

提出对抗训练框架,让模型在被污染奖励下仍能保持强决策能力。

Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks

  • 用对抗训练让模型和攻击者同步优化,提升鲁棒性。
  • 在带污染的随机任务中表现优于现有抗干扰算法。
  • 适用于复杂环境,适合开发抗干扰强化学习系统。

我们研究了上下文内强化学习(ICRL)在奖励污染攻击下的鲁棒性,聚焦于决策预训练变换器(DPT, Lee et al., 2023)。为应对针对DPT的奖励污染攻击,我们提出一种新型对抗训练框架——对抗训练DPT(AT-DPT)。该方法同时训练一组攻击者以污染环境奖励来最小化DPT的真实奖励,并训练DPT模型从污染数据中推断最优动作。我们在标准的多臂赌博机算法上评估该方法,包括专为处理奖励污染设计的稳健基线。结果表明,在学习到的攻击者下,AT-DPT在赌博机设置中显著优于现有方法,并能推广至更复杂的环境,如自适应攻击者和马尔可夫决策过程(MDPs)。AT-DPT展现出作为元强化学习方法,在构建有效抗污染算法方面具有潜力。

原文摘要 · Abstract (English)

We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of reward poisoning attacks targeting the DPT, we propose a novel adversarial training framework, called Adversarially Trained DPT (AT-DPT). Our method simultaneously trains a population of attackers to minimize the true reward of the DPT by poisoning environment rewards, and a DPT model to infer optimal actions from the poisoned data. We evaluate the effectiveness of our approach against standard bandit algorithms, including robust baselines designed to handle reward contamination. Our results show that AT-DPT significantly outperforms them in bandit settings under a learned attacker, and generalizes to more complex environments such as adaptive attackers and MDPs. It shows promise in ICRL as a meta-RL approach to learning effective corruption-robust algorithms.

强化学习对抗训练鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。