arXiv:2607.18496cs.CRcs.AI2026-07

用权威机构信息自动检测大模型在安全知识上的漏洞

Towards an Automated Test of LLM Security Knowledge

论文配图:Towards an Automated Test of LLM Security Knowledge
图 1 · 摘自论文原文
  • 利用消费者保护机构资料识别大模型回答不一致处
  • 在身份盗窃和冒充诈骗上区分出有无足够安全知识的模型
  • 适合安全评测与大模型可靠性研究者使用

大型语言模型(LLMs)被广泛应用于软件、硬件及人机安全任务。因此,评估其在安全任务中的表现成为研究热点,尤其关注模型安全知识是否存在不足。现有方法多依赖人工构建挑战问题或基准测试,需大量人力与安全专业背景。本文提出一种部分自动化的方法,通过权威消费者保护机构(CPAs)提供的信息,识别大模型回答中的不稳定性,从而揭示知识缺口。我们在身份盗窃和冒充诈骗两个安全主题上,对5个主流大模型(来自Gemini和GPT两大系列)进行了验证,基于6个机构公开的关于身份盗窃与冒充诈骗的信息。结果表明,该方法能够有效区分具备充分知识与缺乏知识的模型。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security "knowledge" may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and security expertise to design and execute. We introduce a partially-automated method for assessing LLM knowledge of a security area. The method uses authoritative information from Consumer Protection Agencies (CPAs) to identify instability in LLM responses that can be indicative of knowledge gaps. We demonstrate the method for 2 security topics, identity theft and impostor scams, and 5 LLMs in 2 leading LLM families, Gemini and GPT, using publicly available information about identity theft and impostor scams from 6 CPAs. The method distinguishes between models that have and don't have sufficient knowledge to accurately identify the security topics in text narratives.

大模型安全评测方法知识检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。