arXiv:2501.01872cs.CL2025-01EMNLP被引 1

用对比提问诱导大模型输出有害内容,揭示其推理漏洞。

Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions

  • 通过构造语义对立的提问和对抗模板,诱导模型产生越狱响应。
  • 在六类模型上实现约44%攻击成功率,显著高于现有方法。
  • 提出双反向推理防御机制,可识别恶意意图并拒绝有害输出。

大型语言模型尽管经过与人类价值观和伦理原则的对齐训练,仍易受复杂越狱攻击影响,这些攻击利用其推理能力进行隐蔽诱导。现有安全措施多能检测明显恶意意图,却难以应对由推理驱动的细微漏洞。本文提出POATE(极性相反查询生成、对抗模板构建与扩展),一种新型越狱技术,通过对比推理激发模型产生不道德响应。该方法构建语义相反的意图,并结合对抗性模板,以极强隐蔽性引导模型输出有害内容。我们在六种不同参数规模的语言模型家族中进行了广泛评估,验证了该攻击的鲁棒性,在多种场景下达到约44%的攻击成功率,显著优于现有方法。为应对这一威胁,我们进一步提出意图感知思维链(Intent-Aware CoT)与逆向思维链(Reverse Thinking CoT),通过分解查询识别恶意意图,并反向推理评估与拒绝潜在有害响应。这两种方法有效提升模型推理鲁棒性,增强对对抗性攻击的防御能力。

原文摘要 · Abstract (English)

Large language models, despite extensive alignment with human values and ethical principles, remain vulnerable to sophisticated jailbreak attacks that exploit their reasoning abilities. Existing safety measures often detect overt malicious intent but fail to address subtle, reasoning-driven vulnerabilities. In this work, we introduce POATE (Polar Opposite query generation, Adversarial Template construction, and Elaboration), a novel jailbreak technique that harnesses contrastive reasoning to provoke unethical responses. POATE crafts semantically opposing intents and integrates them with adversarial templates, steering models toward harmful outputs with remarkable subtlety. We conduct extensive evaluation across six diverse language model families of varying parameter sizes to demonstrate the robustness of the attack, achieving significantly higher attack success rates (~44%) compared to existing methods. To counter this, we propose Intent-Aware CoT and Reverse Thinking CoT, which decompose queries to detect malicious intent and reason in reverse to evaluate and reject harmful responses. These methods enhance reasoning robustness and strengthen the model's defense against adversarial exploits.

越狱攻击推理安全对抗样本防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。