arXiv:2605.16312cs.LGcs.AI2026-05

攻击者通过移除合法动作破坏自对弈强化学习,显著降低模型性能。

When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

  • 攻击者在自对弈中针对性删除对手的合法动作,影响决策
  • 针对高价值决策点攻击,使模型性能下降超随机基线
  • 适用于多种强化学习算法,且攻击效果可迁移、难恢复

我们研究自对弈强化学习中的对抗性动作屏蔽:攻击者有选择地从受害者的行为集合中移除合法动作。与观测或动作扰动不同,这种移除在智能体行动前就消除了决策选项。在从6到5,531个信息状态的扑克游戏以及两个非扑克领域中,学习到的屏蔽策略造成的损害远超随机屏蔽和学习扰动基线。该攻击对Q-learning、PPO、NFSP、神经网络NFSP和DQN等多类受害模型均有效,具有跨智能体迁移性,并在自对弈训练中被放大,且在延长遮蔽训练下无法恢复。机制上,攻击者聚焦于高价值决策点,可通过可达加权条件动作容量(CAC$_w$)和价值加权精炼版CAC$_v$进行刻画。这些结果揭示了动作可用性是自对弈强化学习中一个独特的鲁棒性维度。

原文摘要 · Abstract (English)

We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unlike observation or action perturbations, removal eliminates decision options before the agent acts. Across poker games scaling from 6 to 5,531 information states and two non-poker domains, learned masking causes substantially more damage than random masking and learned perturbation baselines. The attack persists across Q-learning, PPO, NFSP, neural NFSP, and DQN victims; transfers across agents; is amplified by self-play; and shows no recovery under extended masked training. Mechanistically, the adversary targets high-value decision points, captured by reach-weighted contingent action capacity (CAC$_w$) and a value-weighted refinement CAC$_v$. These results identify action availability as a distinct robustness surface in self-play RL.

对抗攻击强化学习自对弈动作屏蔽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。