arXiv:2507.22564cs.CLcs.AI2025-07AAAI被引 2

利用认知偏见协同效应,突破大模型安全防护

Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs

  • 设计框架结合多种认知偏见生成对抗提示
  • 在30个模型上实现60.1%攻击成功率,远超现有方法
  • 为提升模型安全性提供认知科学新视角,适合安全研究者

大型语言模型在多项任务中表现卓越,但其安全机制仍易受利用认知偏见的攻击。与以往聚焦提示工程或算法操控的方法不同,本文揭示了多偏见协同作用对模型防护的破坏力。提出CognitiveAttack框架,通过监督微调与强化学习,生成融合优化偏见组合的提示,有效绕过安全策略,同时保持高攻击成功率。实验显示,该方法在30种不同LLM中均暴露显著漏洞,尤其在开源模型中更为明显,攻击成功率达60.1%,远高于当前最优黑盒方法PAP(31.6%),暴露出现有防御机制的关键缺陷。研究将认知科学与模型安全结合,开辟新研究路径,助力构建更鲁棒、更符合人类对齐的AI系统。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate impressive capabilities across a wide range of tasks, yet their safety mechanisms remain susceptible to adversarial attacks that exploit cognitive biases -- systematic deviations from rational judgment. Unlike prior jailbreaking approaches focused on prompt engineering or algorithmic manipulation, this work highlights the overlooked power of multi-bias interactions in undermining LLM safeguards. We propose CognitiveAttack, a novel red-teaming framework that systematically leverages both individual and combined cognitive biases. By integrating supervised fine-tuning and reinforcement learning, CognitiveAttack generates prompts that embed optimized bias combinations, effectively bypassing safety protocols while maintaining high attack success rates. Experimental results reveal significant vulnerabilities across 30 diverse LLMs, particularly in open-source models. CognitiveAttack achieves a substantially higher attack success rate compared to the SOTA black-box method PAP (60.1% vs. 31.6%), exposing critical limitations in current defense mechanisms. These findings highlight multi-bias interactions as a powerful yet underexplored attack vector. This work introduces a novel interdisciplinary perspective by bridging cognitive science and LLM safety, paving the way for more robust and human-aligned AI systems.

模型安全认知偏见对抗攻击红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。