针对离线强化学习的数据投毒攻击,按样本影响力分配扰动预算。
Optimal Perturbation Budget Allocation for Data Poisoning in Offline Reinforcement Learning
- 根据时序差分误差敏感度,全局优化扰动分配策略。
- 在D4RL数据集上实现最高80%性能下降,扰动量极小且隐蔽。
- 适合研究模型安全与对抗攻击防御的科研人员。
离线强化学习(Offline RL)虽能从静态数据集中优化策略,但易受数据投毒攻击。现有攻击方法多采用局部均匀扰动,对所有样本一视同仁,导致扰动预算浪费于低影响样本,且因显著统计偏差而缺乏隐蔽性。本文提出一种新型全局扰动预算分配攻击策略。基于理论洞察:样本对价值函数收敛的影响与其时序差分(TD)误差成正比,将攻击建模为全局资源分配问题。推导出闭式解,在全局L2约束下,扰动幅度与TD误差敏感度成比例分配。在D4RL基准上的实证结果表明,该方法显著优于基线策略,在仅使用极小扰动的情况下,实现高达80%的性能下降,且可规避主流统计与谱检测防御机制。
原文摘要 · Abstract (English)
Offline Reinforcement Learning (RL) enables policy optimization from static datasets but is inherently vulnerable to data poisoning attacks. Existing attack strategies typically rely on locally uniform perturbations, which treat all samples indiscriminately. This approach is inefficient, as it wastes the perturbation budget on low-impact samples, and lacks stealthiness due to significant statistical deviations. In this paper, we propose a novel Global Budget Allocation attack strategy. Leveraging the theoretical insight that a sample's influence on value function convergence is proportional to its Temporal Difference (TD) error, we formulate the attack as a global resource allocation problem. We derive a closed-form solution where perturbation magnitudes are assigned proportional to the TD-error sensitivity under a global L2 constraint. Empirical results on D4RL benchmarks demonstrate that our method significantly outperforms baseline strategies, achieving up to 80% performance degradation with minimal perturbations that evade detection by state-of-the-art statistical and spectral defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。