提出一种隐蔽的分布式攻击,让单步检测失效但能被时间相关性识别。
The Distributed Detectability Band Against Marginal-Preserving Attacks
- 用高斯耦合自回归构造攻击,保持每步评分分布不变以逃避检测。
- 在3.0危害水平下,攻击与正常行为的差异仅0.013,几乎无法察觉。
- 基于时间相关性的监控可检测攻击,适合防范隐蔽协同破坏。
AI控制监控器通过评分单个智能体动作来检测异常行为,但真实伤害可能分散在多个看似正常的步骤中,每个步骤单独均低于报警阈值。本文构建了一种保边缘、编码相关性的分布式破坏攻击,采用高斯耦合自回归AR(1)结构:每步监控评分的边缘分布严格等于正常情况,因此均值、最大值、顶部k尾部、阈值监控器(监测器A)被构造上击败;而伤害则编码在时间相关性结构中。论文围绕三个审稿人要求的关卡展开:(1) 可实现性关卡:该隐蔽攻击在所有测试危害水平(最高3.0)下,与正常行为的K-S距离仅为0.013(近乎零),证实伤害可完全脱离单步边缘分布,且可实现性不受危害限制。(2) 监控器A与B的协调性:我们形式化证明,该针对监控器A设计的攻击,在另一类评分监控器B(如序列型:CUSUM、SPRT、HMM-LR、游程检验、自相关、滑动窗口逻辑回归)下仍保持边缘不变,并涵盖那些具有时间特征的评分函数的最坏情况。(3) 非空可检测区间:监控器A在1%误报率目标下仅达AUC 0.52(随机水平);而监控器B在相同条件下达到AUC 0.79–0.97,随着危害分散在更多步骤中,监控器A退化至随机水平,而监控器B稳定维持在AUC ~0.95。这些结果证明存在非空可检测区间,并刻画了亚阈值破坏的边界:分布形状监控器因构造而失效;时间相关性监控器可检测但非最优。
原文摘要 · Abstract (English)
AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm. We construct a marginal-preserving, correlation-encoded distributed-sabotage attack using a Gaussian-copula AR(1) construction: the per-step monitor-score marginal is held exactly equal to benign, so mean, max, top-k tail, and threshold monitors (Monitor A) are defeated by construction, while harm is encoded in the temporal correlation structure. We sequence the paper around three reviewer-mandated gates. (1) Realizability gate: the stealthy attack achieves KS-distance to benign of 0.013 (effectively zero) at all tested harm levels up to 3.0, confirming that harm is fully decoupled from the per-step marginal and realizability is not harm-limited. (2) Monitor-A-vs-B reconciliation: we show formally that the attack, built against Monitor A's score marginal, remains marginal-preserving under a different-score Monitor B (the correlation/sequence family: CUSUM, SPRT, HMM-LR, runs test, autocorrelation, windowed logistic), and scope worst-case claims to score functions that admit a temporal signature. (3) Non-empty detectability band: Monitor A achieves AUC 0.52 (chance); Monitor B spans AUC 0.79-0.97 at the same 1% FPR target, and as harm is amortized over more steps Monitor A collapses to chance while Monitor B holds at AUC ~0.95. These results demonstrate a non-empty detectability band and characterize the sub-threshold sabotage frontier: distribution-shape monitors fail by construction; temporal-correlation monitors can detect but are not trivially optimal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。