破解模型窃取防御的单客户端假设,揭示协同攻击威胁。
AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses

- 构建开源框架CerberusAI,模拟分布式攻击场景。
- 轮询查询策略可绕过PRADA等主流防御,检测率大幅下降。
- 适合军事与关键基础设施安全研究者参考。
保障部署于军事指挥控制(C2)系统和关键基础设施中的人工智能模型安全,对维持信息优势至关重要。模型提取攻击(MEAs)威胁显著,可使对手复制专有模型、泄露敏感信息并提前准备离线对抗攻击。然而,现有防御多依赖单客户端假设(SCA),即默认攻击来自孤立身份。本文系统证明,在高级持续性威胁(APTs)等协同攻击者存在下,该假设根本无效。我们提出开源框架CerberusAI,用于可复现的模型窃取研究,并模拟分布式攻击。实证表明,如PRADA等成熟防御机制可被基础轮询查询策略绕过,导致检测性能显著下降。此外,全局聚合方法亦可通过自适应流量混杂被彻底失效。结果强调需转向状态化、身份无关的防御架构。本论文原发表于2026年5月12-13日英国巴斯举行的国际军事通信与信息系统会议(ICMCIS),获最佳论文奖。
原文摘要 · Abstract (English)
Ensuring the protection of Artificial Intelligence (AI) models deployed in military Command and Control (C2) systems and critical infrastructure is essential for maintaining information superiority. Model Extraction Attacks (MEAs) pose a significant threat, as they enable adversaries to replicate proprietary models, compromise protected information, and prepare offline adversarial attacks. However, current defense strategies predominantly rely on the Single Client Assumption (SCA), which is the implicit assumption that attacks originate from isolated identities. This work systematically demonstrates that the SCA is fundamentally invalid in the presence of coordinated threat actors, such as Advanced Persistent Threats (APTs). We introduce a modular, open-source framework called CerberusAI for reproducible model-stealing research, and use it to simulate distributed attack scenarios. Our empirical evaluation shows that well-established defense mechanisms, such as Protecting Against Deep Neural Network Model Stealing Attacks (PRADA), can be bypassed by basic round-robin query distribution strategies, resulting in a significant reduction in detection performance. Furthermore, we demonstrate that even global aggregation approaches can be rendered operationally useless through adaptive traffic mixing. These results highlight the need for a paradigm shift towards stateful, identity-independent defense architectures in the field of model extraction attacks. This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026 and won the best paper award.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。