首次精确界定线性MDP中奖励投毒的可行边界。
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs

- 构建理论框架,精准判断何时可成功投毒
- 揭示部分模型天生抗攻击,无需高成本即可识别
- 适用于深度强化学习环境,兼具理论与实践价值
我们研究强化学习中的奖励投毒攻击,即攻击者在有限预算内操纵奖励,迫使目标强化学习智能体采用符合攻击者目标的策略。以往研究多关注攻击成功的充分条件,仅有少数探讨攻击不可行的情形。本文首次对线性马尔可夫决策过程(Linear MDP)下的奖励投毒攻击性提供了必要性与充分性的精确刻画。该刻画清晰区分了易受攻击的强化学习实例与内在鲁棒的实例——后者即使使用原始非鲁棒的强化学习算法,也无法在不付出巨大代价的情况下被攻击。我们的理论不仅限于线性MDP,通过将深度强化学习环境近似为线性MDP,证明该框架能有效区分攻击可行性并高效攻击脆弱实例,展示了其理论与实际意义。
原文摘要 · Abstract (English)
We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards within constrained budgets to force the target RL agent to adopt a policy that aligns with the attacker's objectives. Prior works on reward poisoning mainly focused on sufficient conditions to design a successful attacker, while only a few studies discussed the infeasibility of targeted attacks. This paper provides the first precise necessity and sufficiency characterization of the attackability of a linear MDP under reward poisoning attacks. Our characterization draws a bright line between the vulnerable RL instances, and the intrinsically robust ones which cannot be attacked without large costs even running vanilla non-robust RL algorithms. Our theory extends beyond linear MDPs -- by approximating deep RL environments as linear MDPs, we show that our theoretical framework effectively distinguishes the attackability and efficiently attacks the vulnerable ones, demonstrating both the theoretical and practical significance of our characterization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。