arXiv:2512.21238cs.SEcs.CR2025-12被引 1

评测五大模型的软件安全理解能力,发现其高阶推理能力严重不足

Assessing the Software Security Comprehension of Large Language Models

  • 用布鲁姆认知框架分六层评估模型安全理解力
  • 低阶任务准确率超80%,高阶任务下降至不足50%
  • 发现51种常见误解模式,适合安全开发与评估研究者参考

大型语言模型在软件开发中应用日益广泛,但其软件安全理解水平尚不明确。本研究系统评估了五款主流LLM(GPT-4o-Mini、GPT-5-Mini、Gemini-2.5-Flash、Llama-3.1、Qwen-2.5)在软件安全领域的认知能力,采用布鲁姆认知分类法作为评估框架,涵盖记忆、理解、应用、分析、评价和创造六个维度。方法整合了多种数据集:精选多选题、漏洞代码片段(SALLM)、软件安全导论课程测验、真实案例研究(XBOW)以及安全软件工程课程的项目创作任务。结果显示,尽管模型在事实回忆和已知漏洞识别等低阶任务上表现良好(准确率>80%),但在需要推理、架构评估和安全系统设计的高阶任务中性能显著下降(平均准确率<50%)。除报告总体准确率外,本文提出‘软件安全知识边界’概念,界定模型可稳定可靠执行的最高认知层级。此外,识别出51种跨认知层次的常见误解模式。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in software development, but their level of software security expertise remains unclear. This work systematically evaluates the security comprehension of five leading LLMs: GPT-4o-Mini, GPT-5-Mini, Gemini-2.5-Flash, Llama-3.1, and Qwen-2.5, using Blooms Taxonomy as a framework. We assess six cognitive dimensions: remembering, understanding, applying, analyzing, evaluating, and creating. Our methodology integrates diverse datasets, including curated multiple-choice questions, vulnerable code snippets (SALLM), course assessments from an Introduction to Software Security course, real-world case studies (XBOW), and project-based creation tasks from a Secure Software Engineering course. Results show that while LLMs perform well on lower-level cognitive tasks such as recalling facts and identifying known vulnerabilities, their performance degrades significantly on higher-order tasks that require reasoning, architectural evaluation, and secure system creation. Beyond reporting aggregate accuracy, we introduce a software security knowledge boundary that identifies the highest cognitive level at which a model consistently maintains reliable performance. In addition, we identify 51 recurring misconception patterns exhibited by LLMs across Blooms levels.

大模型测评软件安全认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。