arXiv:2506.10022cs.CRcs.AI2025-06ACL被引 14

构建首个恶意代码生成越狱攻击基准,揭示大模型安全短板

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges

  • 设计包含3520条越狱提示的MalwareBench数据集
  • 主流大模型对恶意请求平均拒绝率仅60.93%,多方法组合下降至39.92%
  • 适合关注大模型安全、代码生成风险的研究者与开发者

大型语言模型(LLMs)的广泛应用引发了对其安全性的担忧,尤其是其易受精心设计提示引发的越狱攻击影响,导致生成恶意输出。尽管已有研究关注大模型的一般安全能力,但其在代码生成场景下的越狱攻击脆弱性仍缺乏系统探索。为此,我们提出MalwareBench,一个包含3520个恶意代码生成越狱提示的基准数据集,用于评估大模型对此类威胁的鲁棒性。该数据集基于320个手工构造的恶意代码生成需求,覆盖11种越狱方法和29类代码功能。实验表明,主流大模型对恶意代码生成请求的拒绝对抗能力有限,平均拒绝率为60.93%;当结合多种越狱算法时,该比率下降至39.92%。本工作凸显了当前大模型在代码安全方面仍面临严峻挑战。

原文摘要 · Abstract (English)

The widespread adoption of Large Language Models (LLMs) has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs. While prior research has been conducted on general security capabilities of LLMs, their specific susceptibility to jailbreak attacks in code generation remains largely unexplored. To fill this gap, we propose MalwareBench, a benchmark dataset containing 3,520 jailbreaking prompts for malicious code-generation, designed to evaluate LLM robustness against such threats. MalwareBench is based on 320 manually crafted malicious code generation requirements, covering 11 jailbreak methods and 29 code functionality categories. Experiments show that mainstream LLMs exhibit limited ability to reject malicious code-generation requirements, and the combination of multiple jailbreak methods further reduces the model's security capabilities: specifically, the average rejection rate for malicious content is 60.93%, dropping to 39.92% when combined with jailbreak attack algorithms. Our work highlights that the code security capabilities of LLMs still pose significant challenges.

大模型安全越狱攻击代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。