评测大模型对真实云架构的威胁建模能力,发现GPT-4.1和Gemini表现最优。
ACSE-Eval: Can LLMs threat model real-world cloud infrastructure?
- 构建100个真实AWS部署场景,含架构、代码和漏洞数据
- Gemini 2.5 Pro在零样本下威胁识别最佳,GPT 4.1少样本更优
- 开源数据集与评估方法,推动自动化安全分析研究
尽管大语言模型在网络安全应用中展现出潜力,但其在真实云部署中识别安全威胁的有效性仍未知。本文提出ACSE-Eval——一个用于评估大模型云安全威胁建模能力的新数据集。该数据集包含100个生产级AWS部署场景,每个场景均配有详细的架构说明、基础设施即代码实现、已知安全漏洞及关联的威胁建模参数。本研究系统评估了大模型在识别安全风险、分析攻击路径和提出缓解策略方面的能力。实验表明,GPT 4.1和Gemini 2.5 Pro在威胁识别上表现优异,其中Gemini 2.5 Pro在零样本(0-shot)条件下表现最佳,而GPT 4.1在少量示例(few-shot)设置中更胜一筹。虽然GPT 4.1整体略占优势,但Claude 3.7 Sonnet生成的威胁模型语义最丰富,却在威胁分类和泛化能力上存在不足。为促进可复现性和推动自动化安全分析研究,我们开源了数据集、评估指标与方法论。
原文摘要 · Abstract (English)
While Large Language Models have shown promise in cybersecurity applications, their effectiveness in identifying security threats within cloud deployments remains unexplored. This paper introduces AWS Cloud Security Engineering Eval, a novel dataset for evaluating LLMs cloud security threat modeling capabilities. ACSE-Eval contains 100 production grade AWS deployment scenarios, each featuring detailed architectural specifications, Infrastructure as Code implementations, documented security vulnerabilities, and associated threat modeling parameters. Our dataset enables systemic assessment of LLMs abilities to identify security risks, analyze attack vectors, and propose mitigation strategies in cloud environments. Our evaluations on ACSE-Eval demonstrate that GPT 4.1 and Gemini 2.5 Pro excel at threat identification, with Gemini 2.5 Pro performing optimally in 0-shot scenarios and GPT 4.1 showing superior results in few-shot settings. While GPT 4.1 maintains a slight overall performance advantage, Claude 3.7 Sonnet generates the most semantically sophisticated threat models but struggles with threat categorization and generalization. To promote reproducibility and advance research in automated cybersecurity threat analysis, we open-source our dataset, evaluation metrics, and methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。