arXiv:2502.19820cs.CLcs.AI2025-02EMNLP被引 52

利用心理渐进原理,用多轮对话诱骗大模型输出有害内容

Foot-In-The-Door: A Multi-turn Jailbreak for LLMs

  • 通过逐步引导的对话设计,让模型自我松懈安全防线
  • 在7个主流模型上平均攻击成功率高达94%
  • 揭示多轮交互中模型自腐蚀风险,适合安全研究者参考

随着大语言模型在现实应用中日益普及,AI安全问题愈发重要。其中,越狱攻击——即通过对抗性提示绕过内置防护机制,诱导模型生成有害内容——成为关键挑战。受心理学‘脚踏进门’效应启发,我们提出FITD,一种新型多轮越狱方法。该方法通过中间桥梁提示逐步提升用户请求的恶意程度,并利用模型自身响应进行对齐,从而诱导出有毒输出。在两个越狱基准上的大量实验表明,FITD在七个广泛使用的模型上平均攻击成功率高达94%,优于现有最先进方法。此外,我们深入分析了大模型自我腐蚀现象,揭示当前对齐策略的脆弱性,强调多轮交互中固有的风险。代码已开源。

原文摘要 · Abstract (English)

Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs. Inspired by psychological foot-in-the-door principles, we introduce FITD,a novel multi-turn jailbreak method that leverages the phenomenon where minor initial commitments lower resistance to more significant or more unethical transgressions. Our approach progressively escalates the malicious intent of user queries through intermediate bridge prompts and aligns the model's response by itself to induce toxic responses. Extensive experimental results on two jailbreak benchmarks demonstrate that FITD achieves an average attack success rate of 94% across seven widely used models, outperforming existing state-of-the-art methods. Additionally, we provide an in-depth analysis of LLM self-corruption, highlighting vulnerabilities in current alignment strategies and emphasizing the risks inherent in multi-turn interactions. The code is available at https://github.com/Jinxiaolong1129/Foot-in-the-door-Jailbreak.

越狱攻击多轮对话模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。