arXiv:2509.05471cs.CRcs.AI2025-09被引 3

测试大模型对伪装攻击的防御能力,发现现有安全机制形同虚设。

Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models

论文配图:Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
图 1 · 摘自论文原文
  • 构建500条伪装攻击提示数据集,含400条恶意与100条良性样本。
  • 模型在伪装攻击下安全评分下降超60%,内容质量严重下滑。
  • 提出七维评估框架,适合安全研究人员与模型开发者使用。

大型语言模型(LLMs)正面临一种新型隐蔽攻击——伪装越狱攻击,其将恶意意图隐藏于看似无害的语言中,绕过现有安全机制。与明显攻击不同,这类攻击利用语言的上下文模糊性和灵活性,使检测变得困难。本文提出一个全新的基准数据集:Camouflaged Jailbreak Prompts,包含500条精心设计的提示(400条有害,100条良性),用于严格测试大模型的安全协议。同时,设计了一个多维度评估框架,从七个维度衡量危害性:安全意识、技术可行性、实施防护、潜在危害、教育价值、内容质量与合规得分。实验结果显示,模型在良性输入下表现良好,但在伪装攻击面前安全性能显著下降,内容质量与合规得分平均降低超过60%。这一巨大差异揭示了当前防御体系的深层漏洞,亟需更精细、自适应的安全策略以保障大模型在真实场景中的可靠部署。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly vulnerable to a sophisticated form of adversarial prompting known as camouflaged jailbreaking. This method embeds malicious intent within seemingly benign language to evade existing safety mechanisms. Unlike overt attacks, these subtle prompts exploit contextual ambiguity and the flexible nature of language, posing significant challenges to current defense systems. This paper investigates the construction and impact of camouflaged jailbreak prompts, emphasizing their deceptive characteristics and the limitations of traditional keyword-based detection methods. We introduce a novel benchmark dataset, Camouflaged Jailbreak Prompts, containing 500 curated examples (400 harmful and 100 benign prompts) designed to rigorously stress-test LLM safety protocols. In addition, we propose a multi-faceted evaluation framework that measures harmfulness across seven dimensions: Safety Awareness, Technical Feasibility, Implementation Safeguards, Harmful Potential, Educational Value, Content Quality, and Compliance Score. Our findings reveal a stark contrast in LLM behavior: while models demonstrate high safety and content quality with benign inputs, they exhibit a significant decline in performance and safety when confronted with camouflaged jailbreak attempts. This disparity underscores a pervasive vulnerability, highlighting the urgent need for more nuanced and adaptive security strategies to ensure the responsible and robust deployment of LLMs in real-world applications.

模型安全越狱攻击评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。