测试开源大模型提示词攻击漏洞,发现近九成可被诱导生成不当内容。
From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
- 提出新评估指标ASP,捕捉模型响应的不确定性与矛盾行为。
- 90%的开源模型在诱导攻击下产生不当输出,部分攻击突破所有14个模型。
- 知名度中等的模型更易受攻击,适合安全研究人员与开发者参考。
近期研究显示大型语言模型易受生成有害或敏感内容的攻击。随着开源大模型在金融、法律、医疗等高影响力领域广泛应用,系统性评估其安全风险对构建可信大模型时代至关重要。本文全面研究了针对14个主流开源模型和3个闭源模型的提示注入攻击,在五个攻击基准上展开实验。现有评估指标多仅关注攻击成功率,忽略了模型响应中的不确定性。为此,本文提出攻击成功率概率(ASP),可捕捉模型可能先拒绝有害请求但后续提供有害引导,或反之的不一致行为,反映攻击可行性中的模糊性。通过系统分析,本文提出一种简单有效的催眠式攻击,使包括Stablelm2、Mistral、Openchat和Vicuna在内的对齐模型产生不当行为,达到约90%的ASP。结果还表明,忽略前缀攻击可使所有14个开源模型失效,在多类别数据集上达到60%以上的ASP。研究发现,知名度适中的模型对提示注入攻击更为脆弱,凸显提升公众意识并优先部署高效缓解策略的必要性。
原文摘要 · Abstract (English)
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically investigating their security risks is becoming increasingly important towards trustworthy LLM era. This paper comprehensively studies effective prompt injection attacks against 14 widely used open-source and three closed-source LLMs on five attack benchmarks. Moreover, existing evaluation metrics mostly only consider the attack success rate, overlooking uncertainty in model responses. Our proposed Attack Success Probability (ASP) additionally captures uncertain behaviors for evaluation, where the model may initially refuse a harmful request but subsequently provide harmful guidance or vice versa, reflecting inconsistency and ambiguity in attack feasibility. By systematically analyzing the effectiveness of prompt injection attacks, we propose a straightforward and effective hypnotism attack; results show that this attack causes aligned language models, including Stablelm2, Mistral, Openchat, and Vicuna, to generate objectionable behaviors, achieving around 90% ASP. They also indicate that ignore prefix attacks can break all 14 open-source LLMs, achieving over 60% ASP on a multi-categorical dataset. We find that moderately well-known LLMs exhibit higher vulnerability to prompt injection attacks, highlighting the need to raise public awareness and prioritize efficient mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。