用心理效应攻击大模型,100%成功率暴露内容生成漏洞
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
- 模仿心理效应设计新型攻击,绕过模型安全机制
- 开源模型攻击成功率100%,闭源模型达95%以上
- 揭示大模型在关键应用中的潜在社会风险,适合安全研究者参考
大型语言模型(LLMs)虽已广泛应用于各行业,但其生成有害内容的敏感性带来严重社会风险。本文设计并测试了新型攻击策略,针对主流LLMs暴露其生成不当内容的脆弱性。这些策略受心理学现象如‘启动效应’、‘安全注意力转移’和‘认知失调’启发,有效突破模型的防御机制。实验显示,在Meta的Llama-3.2、Google的Gemma-2、Mistral的Mistral-NeMo、Falcon的Falcon-mamba、Apple的DCLM、Microsoft的Phi3以及Qwen的Qwen2.5等开源模型上,攻击成功率达100%;在OpenAI的GPT-4o、Google的Gemini-1.5和Claude-3.5等闭源模型上,于AdvBench数据集上的攻击成功率均不低于95%,达到当前最先进水平。该研究强调亟需重新评估生成模型在关键场景中的使用,以降低潜在的负面影响。
原文摘要 · Abstract (English)
Large language models (LLMs) have significantly influenced various industries but suffer from a critical flaw, the potential sensitivity of generating harmful content, which poses severe societal risks. We developed and tested novel attack strategies on popular LLMs to expose their vulnerabilities in generating inappropriate content. These strategies, inspired by psychological phenomena such as the "Priming Effect", "Safe Attention Shift", and "Cognitive Dissonance", effectively attack the models' guarding mechanisms. Our experiments achieved an attack success rate (ASR) of 100% on various open-source models, including Meta's Llama-3.2, Google's Gemma-2, Mistral's Mistral-NeMo, Falcon's Falcon-mamba, Apple's DCLM, Microsoft's Phi3, and Qwen's Qwen2.5, among others. Similarly, for closed-source models such as OpenAI's GPT-4o, Google's Gemini-1.5, and Claude-3.5, we observed an ASR of at least 95% on the AdvBench dataset, which represents the current state-of-the-art. This study underscores the urgent need to reassess the use of generative models in critical applications to mitigate potential adverse societal impacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。