arXiv:2411.13207cs.CRcs.AI2024-11被引 1

评估大模型信息安全意识,发现多数模型存在安全认知短板。

LISAA: A Framework for Large Language Model Information Security Awareness Assessment

  • 构建100个真实场景测试大模型的安全知识与行为反应
  • 多数主流模型信息安全意识仅达中低水平,小模型风险更高
  • 新模型虽有进步,但安全漏洞仍存,适合安全研究人员使用

大型语言模型(LLMs)日益普及,基于其的助手已无处不在。信息安全意识(ISA)是大模型安全中重要却未被充分研究的领域,涵盖模型的安全知识、态度与行为,对理解隐含安全情境及拒绝可能导致用户意外失败的不安全请求至关重要。本文提出LISAA框架,通过自动化方法评估涵盖ISA分类中所有安全主题的100个真实场景。这些场景在隐含安全影响与用户满意度之间制造张力。对领先大模型的应用显示,当前部署普遍存在脆弱性:多数流行模型仅具中等至低水平的ISA,使其用户面临网络安全隐患;部分在网络安全知识基准中表现优异的模型,其ISA排名却相对较低。此外,同一模型家族中的小型版本风险显著更高。尽管新版本模型展现出明显改进,其ISA仍存在显著差距,表明仍有提升空间。我们发布了在线工具以实现新模型的评估。

原文摘要 · Abstract (English)

The popularity of large language models (LLMs) continues to grow, and LLM-based assistants have become ubiquitous. Information security awareness (ISA) is an important yet underexplored area of LLM safety. ISA encompasses LLMs' security knowledge, which has been explored in the past, as well as their attitudes and behaviors, which are crucial to LLMs' ability to understand implicit security context and reject unsafe requests that may cause an LLM to unintentionally fail the user. We introduce LISAA, a comprehensive framework to assess LLM ISA. The proposed framework applies an automated measurement method to a comprehensive set of 100 realistic scenarios covering all security topics in an ISA taxonomy. These scenarios create tension between implicit security implications and user satisfaction. Applying our LISAA framework to leading LLMs highlights a widespread vulnerability affecting current deployments: many popular models exhibit only medium to low ISA levels, exposing their users to cybersecurity threats, and models that rank highly in cybersecurity knowledge benchmarks sometimes achieve relatively low ISA ranking. In addition, we found that smaller variants of the same model family are significantly riskier. Furthermore, while newer model versions demonstrated notable improvements, meaningful gaps in their ISA persist, suggesting that there is room for improvement. We release an online tool that implements our framework and enables the evaluation of new models.

大模型安全信息安全评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。