arXiv:2505.20621cs.LGcs.AI2025-05ICLR被引 4

提出多层级认证防御,提升离线强化学习抗数据污染能力

Multi-level Certified Defense Against Poisoning Attacks in Offline Reinforcement Learning

  • 结合差分隐私构建跨状态与累积奖励的双重认证保障
  • 7%数据被污染时性能下降不超过50%,优于先前0.008%的上限
  • 适用于连续/离散、随机/确定性环境,适合高安全需求场景

与其它机器学习框架类似,离线强化学习因依赖外部数据集而易受投毒攻击,其序列特性更放大了这一风险。为缓解该问题,本文将认证防御扩展至多层级,确保对每状态动作及整体预期累计奖励的鲁棒性。方法基于差分隐私特性,适用于连续与离散空间,以及随机与确定性环境,显著拓展了可实现保证的范围与适用性。实验表明,在训练数据最多7%被污染时,性能下降不超过50%,较之前工作(0.008%)大幅提升,且认证半径扩大5倍。这表明本框架在提升离线强化学习安全性与可靠性方面具有潜力。

原文摘要 · Abstract (English)

Similar to other machine learning frameworks, Offline Reinforcement Learning (RL) is shown to be vulnerable to poisoning attacks, due to its reliance on externally sourced datasets, a vulnerability that is exacerbated by its sequential nature. To mitigate the risks posed by RL poisoning, we extend certified defenses to provide larger guarantees against adversarial manipulation, ensuring robustness for both per-state actions, and the overall expected cumulative reward. Our approach leverages properties of Differential Privacy, in a manner that allows this work to span both continuous and discrete spaces, as well as stochastic and deterministic environments -- significantly expanding the scope and applicability of achievable guarantees. Empirical evaluations demonstrate that our approach ensures the performance drops to no more than $50\%$ with up to $7\%$ of the training data poisoned, significantly improving over the $0.008\%$ in prior work~\citep{wu_copa_2022}, while producing certified radii that is $5$ times larger as well. This highlights the potential of our framework to enhance safety and reliability in offline RL.

强化学习投毒攻击认证防御差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。