为生成式AI设计安全评估框架,有效降低攻击成功率。
STRIDE-AI: A Threat Modeling Framework for Generative AI Security Assessment

- 基于STRIDE模型改造,构建六阶段AI安全评估流程
- 实测将大语言模型攻击成功率从80%降至15%
- 适合企业AI安全团队快速落地风险评估
传统网络安全方法针对确定性系统,无法应对AI的随机特性,导致模型逆向、数据污染和提示注入等攻击频发。行业报告显示,多数部署AI的组织缺乏专门安全策略,且对抗攻击逐年上升。本文提出STRIDE-AI框架,连接高阶风险标准(NIST AI RMF)与技术漏洞分类(OWASP LLM Top 10)。该框架包含六阶段评估生命周期,引入适配AI系统的STRIDE威胁建模方法,并通过定制化Web工具实现操作化。初步验证基于黑盒测试一个已部署的LLM聊天机器人,在沙箱案例中将攻击成功率达由80%降至15%。
原文摘要 · Abstract (English)
Traditional cybersecurity methodologies target deterministic systems and fail to address the probabilistic nature of AI, leaving systems vulnerable to attack vectors such as model inversion, data poisoning, and prompt injection. Recent industry reports indicate that a majority of organizations deploying AI lack a dedicated security strategy, with adversarial attacks increasing rapidly year-over-year. We present \textit{STRIDE-AI}, a framework that bridges the gap between high-level risk standards (NIST AI RMF) and technical vulnerability taxonomies (OWASP LLM Top 10). The framework defines a six-phase assessment lifecycle, introduces a threat modeling adaptation of classical STRIDE for AI systems, and is operationalized through a purpose-built web tool. We provide an initial validation of the approach through a black-box assessment of a deployed LLM chatbot, which successfully reduced the attack success rate from 80\% to 15\% in our sandbox case study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。