arXiv:2502.00072cs.CRcs.AI2025-02被引 12

现有LLM安全评估无法反映真实威胁,需结合攻击者行为与实际影响。

LLM Cyber Evaluations Don't Capture Real-World Risk

  • 提出融合威胁行为与影响的综合风险评估框架
  • 前沿模型在真实任务中准确率中等但合规率高
  • 适合关注真实网络安全风险的研究者与从业者

大型语言模型(LLMs)在网络安全应用中表现日益突出,既带来防御增强潜力,也伴随固有风险。本文指出,当前对LLM安全风险的评估方法与真实世界影响目标脱节。仅衡量模型能力不足以评估风险,必须结合攻击者采纳行为和潜在影响进行综合分析。我们提出了一个针对LLM网络安全能力的风险评估框架,并以语言模型作为网络安全助手为例开展案例研究。对前沿模型的评估显示,其合规率较高,但在真实网络安全辅助任务中的准确率仅为中等。然而,该使用场景因缺乏显著操作优势和影响潜力,整体风险被判定为中等。基于此,我们建议加强产学研协作、更真实模拟攻击者行为,并在评估中引入经济指标,以使研究方向更贴近真实世界影响。本工作为有效评估与缓解LLM驱动的网络安全风险迈出关键一步。

原文摘要 · Abstract (English)

Large language models (LLMs) are demonstrating increasing prowess in cybersecurity applications, creating creating inherent risks alongside their potential for strengthening defenses. In this position paper, we argue that current efforts to evaluate risks posed by these capabilities are misaligned with the goal of understanding real-world impact. Evaluating LLM cybersecurity risk requires more than just measuring model capabilities -- it demands a comprehensive risk assessment that incorporates analysis of threat actor adoption behavior and potential for impact. We propose a risk assessment framework for LLM cyber capabilities and apply it to a case study of language models used as cybersecurity assistants. Our evaluation of frontier models reveals high compliance rates but moderate accuracy on realistic cyber assistance tasks. However, our framework suggests that this particular use case presents only moderate risk due to limited operational advantages and impact potential. Based on these findings, we recommend several improvements to align research priorities with real-world impact assessment, including closer academia-industry collaboration, more realistic modeling of attacker behavior, and inclusion of economic metrics in evaluations. This work represents an important step toward more effective assessment and mitigation of LLM-enabled cybersecurity risks.

LLM安全风险评估网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。