首次系统分析状态空间模型的安全风险与认知隐患
Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks
- 构建五层攻击面框架,提出谱敏感性与状态完整性破坏理论
- 发现三类新攻击:谱对抗、延迟触发后门、状态过载,实测攻击效果提升6倍以上
- 针对基因组、医疗、网络安全场景设计攻防对策,适合作为安全研究参考
状态空间模型(SSMs)——包括结构化SSM(S4、S4D、DSS、S5)、选择性SSM(Mamba、Mamba-2)以及混合架构(Jamba)——已被部署于基因组分析、临床时序预测和网络安全日志处理等高风险长上下文应用中。尽管其线性时间复杂度具有吸引力,但其压缩状态递归结构的安全属性尚未被研究。本文首次系统性地揭示了SSMs在安全性、安全性及认知风险方面的威胁。主要贡献包括:(1) 构建形式化威胁框架,引入攻击面分层模型、状态完整性破坏(StIV)、跨上下文放大率$\mathcal{X}_\mathcal{S}$以及基于$H_\infty$范数的谱敏感性命题;(2) 提出三类新型攻击:谱对抗攻击(利用传递函数增益)、延迟触发状态后门(注入后数千步激活)、状态容量饱和(熵洪水导致静默遗忘);(3) 扩展14项MITRE ATLAS技术覆盖完整攻击链;(4) 建立六类攻击者分类与基因组、临床、网络安全领域的杀伤链;(5) 提出四项基于状态压缩机制的认知风险假设;(6) 提出符合CREST、NIST AI 600-1与欧盟人工智能法案的治理缓解方案;(7) 实验验证:靶向基因组注入实现$\mathrm{StIV}=0.519$,较随机注入高6.0倍($p<0.001$);PGD状态注入使输出扰动达随机攻击的156倍;SSD结构提取在$O(N^2)$查询复杂度下完成,相较$O(N^3)$实现$N\times$加速。预训练检查点的验证详见附录。
原文摘要 · Abstract (English)
State-Space Models (SSMs) -- structured SSMs (S4, S4D, DSS, S5), selective SSMs (Mamba, Mamba-2), and hybrid architectures (Jamba) -- are deployed in safety-critical long-context applications: genomic analysis, clinical time-series forecasting, and cybersecurity log processing. Their linear-time scaling is compelling, yet the security properties of their compressed-state recurrent architectures remain unstudied. We present the first systematic treatment of SSM safety, security, and cognitive risks. Seven contributions: (1) Formal threat framework -- SSM Attack Surface (five layers), State Integrity Violation (StIV), Cross-Context Amplification Ratio $\mathcal{X}_\mathcal{S}$, and a Spectral Sensitivity Proposition grounded in the $H_\infty$ norm. (2) Three novel attack classes: spectral adversarial attacks (transfer-function gain exploitation), delayed-trigger stateful backdoors (activate thousands of steps after injection), and state capacity saturation (entropy flooding forces silent forgetting). (3) 14 MITRE ATLAS technique extensions across the full tactic chain. (4) Six-profile attacker taxonomy with kill chains for genomics, clinical, and cybersecurity domains. (5) Four cognitive risk hypotheses grounded in state-compression mechanics. (6) Governance-aligned mitigations mapped to CREST, NIST AI 600-1, and EU AI Act. (7) Empirical evaluation: targeted genomic injection achieves $\mathrm{StIV}=0.519$ vs. $0.086$ random ($6.0\times$, $p<0.001$); PGD state injection achieves $156\times$ output perturbation over random; SSD-structured extraction confirmed at $O(N^2)$ vs. $O(N^3)$ query complexity ($N\times$ speedup). Validation on pretrained checkpoints is detailed in the Appendix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。