arXiv:2410.16222cs.LGcs.AI2024-10被引 11

提出可解释的N元语法困惑度模型,统一评估大模型越狱攻击有效性。

An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks

  • 基于1万亿词的N元语法模型,实现无需依赖LLM的非参数化评估
  • 发现当前越狱攻击成功率低于以往报告,离散优化方法更优
  • 揭示有效攻击多利用罕见或不存在于真实文本中的双词组合

大量越狱攻击被提出以诱导经过安全调优的大语言模型生成有害内容。这些方法在原始设置下大多成功迫使目标输出,但在流畅性和计算开销上差异显著。本文提出一种统一威胁模型,用于原则性比较各类攻击。该模型通过在1万亿词上构建N元语法语言模型,判断给定越狱是否可能出现在文本分布中。与基于模型的困惑度不同,该方法不依赖特定模型、非参数化且天然可解释。我们将其应用于主流攻击并首次实现公平基准测试。经广泛比较发现,对现代安全调优模型的攻击成功率低于先前报道,基于离散优化的攻击显著优于近期基于LLM的攻击。由于其可解释性,本模型支持对越狱攻击的全面分析:有效攻击通常利用真实文本中缺失或极罕见的双词组合,如特定于Reddit或代码数据集的短语。

原文摘要 · Abstract (English)

A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but their attacks vary substantially in fluency and computational effort. In this work, we propose a unified threat model for the principled comparison of these methods. Our threat model checks if a given jailbreak is likely to occur in the distribution of text. For this, we build an N-gram language model on 1T tokens, which, unlike model-based perplexity, allows for an LLM-agnostic, nonparametric, and inherently interpretable evaluation. We adapt popular attacks to this threat model, and, for the first time, benchmark these attacks on equal footing with it. After an extensive comparison, we find attack success rates against safety-tuned modern models to be lower than previously presented and that attacks based on discrete optimization significantly outperform recent LLM-based attacks. Being inherently interpretable, our threat model allows for a comprehensive analysis and comparison of jailbreak attacks. We find that effective attacks exploit and abuse infrequent bigrams, either selecting the ones absent from real-world text or rare ones, e.g., specific to Reddit or code datasets.

大模型安全越狱攻击可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。