arXiv:2506.11172cs.LGcs.AI2025-06

攻击离线强化学习的数据覆盖,仅1%污染就让模型性能下降90%

Collapsing Sequence-Level Data-Policy Coverage via Poisoning Attack in Offline Reinforcement Learning

  • 用序列级覆盖率衡量数据分布,发现误差会指数放大
  • 通过污染罕见决策模式,使数据覆盖骤降导致性能崩塌
  • 揭示离线RL安全漏洞,适合关注AI安全的研究者

离线强化学习严重依赖预收集数据对目标策略分布的覆盖。现有研究致力于提升数据-策略覆盖以缓解分布偏移,却忽视了覆盖不足带来的安全风险,且单步分析与离线RL多步决策本质不一致。为此,我们引入序列级浓缩系数量化覆盖程度,并通过理论分析揭示其对估计误差上界存在指数级放大效应。基于此,提出序列级数据-策略覆盖坍缩(CSDPC) poisoning攻击。考虑到离线RL数据的连续性,将状态-动作对转化为决策单元,提取捕捉多步行为的代表性决策模式,识别可能引发覆盖不足的稀有模式并加以污染,从而降低覆盖度并加剧分布偏移。实验表明,仅污染1%的数据集即可使智能体性能下降90%。该发现为分析和保障离线强化学习的安全性提供了新视角。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) heavily relies on the coverage of pre-collected data over the target policy's distribution. Existing studies aim to improve data-policy coverage to mitigate distributional shifts, but overlook security risks from insufficient coverage, and the single-step analysis is not consistent with the multi-step decision-making nature of offline RL. To address this, we introduce the sequence-level concentrability coefficient to quantify coverage, and reveal its exponential amplification on the upper bound of estimation errors through theoretical analysis. Building on this, we propose the Collapsing Sequence-Level Data-Policy Coverage (CSDPC) poisoning attack. Considering the continuous nature of offline RL data, we convert state-action pairs into decision units, and extract representative decision patterns that capture multi-step behavior. We identify rare patterns likely to cause insufficient coverage, and poison them to reduce coverage and exacerbate distributional shifts. Experiments show that poisoning just 1% of the dataset can degrade agent performance by 90%. This finding provides new perspectives for analyzing and safeguarding the security of offline RL.

离线RL安全攻击数据污染分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。