测试GPT模型在字符级和语义级攻击下的脆弱性,揭示安全机制短板。
Robustness of Large Language Models Against Adversarial Attacks
- 用字符级扰动和越狱提示双重方式测试模型鲁棒性
- 三数据集上模型对攻击响应差异显著,存在明显漏洞
- 适合关注大模型安全与对抗训练的研究者参考
大型语言模型(LLMs)在各类应用中日益普及,亟需对其抗对抗攻击能力进行严格评估。本文针对GPT系列模型开展全面研究,采用两种不同评估方法:第一种在输入提示中引入字符级文本攻击,测试模型在StanfordNLP/IMDB、Yelp Reviews和SST-2三个情感分类数据集上的表现;第二种使用越狱提示挑战模型的安全机制。实验结果表明,这些模型在面对字符级和语义级对抗攻击时表现出显著的鲁棒性差异,暴露出不同程度的脆弱性。研究强调必须改进对抗训练并强化安全机制,以提升大语言模型的整体鲁棒性。
原文摘要 · Abstract (English)
The increasing deployment of Large Language Models (LLMs) in various applications necessitates a rigorous evaluation of their robustness against adversarial attacks. In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。