arXiv:2510.13792cs.LGcs.AI2025-10

用信息论设计无法被破解的对抗攻击,让强化学习模型学不到真实环境信息。

Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach

  • 基于率-失真信息论,随机扰动环境观测以隐藏真实状态转移信息
  • 证明了攻击下智能体的奖励损失存在理论下界,且对主流算法均有效
  • 适用于多种攻击场景,为防御研究提供新视角,适合安全与强化学习研究者

强化学习(RL)在自动驾驶、金融决策和无人机/机器人控制等安全相关领域广泛应用。为提升系统鲁棒性,研究针对RL系统的对抗攻击至关重要。以往工作多关注确定性攻击策略,受害者可通过逆向操作予以应对。本文提出一种可证明‘不可击败’或‘不可抗’的对抗攻击:攻击者采用率-失真信息论方法,随机改变智能体对状态转移核(或其他属性)的观测,使其在训练过程中无法获取真实转移核的任何或有限信息。我们推导出受害者智能体奖励遗憾的信息论下界,并分析该类攻击对当前主流模型基与非模型基算法的影响。同时,该信息论框架也被扩展至状态观测攻击等其他类型对抗攻击。

原文摘要 · Abstract (English)

Reinforcement learning (RL) for the Markov Decision Process (MDP) has emerged in many security-related applications, such as autonomous driving, financial decisions, and drone/robot algorithms. In order to improve the robustness/defense of RL systems against adversaries, studying various adversarial attacks on RL systems is very important. Most previous work considered deterministic adversarial attack strategies in MDP, which the recipient (victim) agent can defeat by reversing the deterministic attacks. In this paper, we propose a provably ``invincible'' or ``uncounterable'' type of adversarial attack on RL. The attackers apply a rate-distortion information-theoretic approach to randomly change agents' observations of the transition kernel (or other properties) so that the agent gains zero or very limited information about the ground-truth kernel (or other properties) during the training. We derive an information-theoretic lower bound on the recipient agent's reward regret and show the impact of rate-distortion attacks on state-of-the-art model-based and model-free algorithms. We also extend this notion of an information-theoretic approach to other types of adversarial attack, such as state observation attacks.

强化学习对抗攻击信息论安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。