用人类认知偏见设计攻击,让大模型说出危险内容
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- 模仿人类思维捷径简化恶意提示
- 在多个主流模型上成功绕过安全机制
- 用排名方法更精准评估内容危害性
尽管大型语言模型(LLMs)表现优异,仍易受越狱攻击,危及安全机制。现有研究多依赖暴力优化或手动设计,难以发现真实场景中的风险。为此,我们提出新攻击框架ICRT,借鉴人类认知中的启发式与偏见。利用简单性效应,通过认知分解降低恶意提示复杂度;同时运用相关性偏见重组提示,增强语义一致性,有效诱导有害输出。此外,引入基于排名的有害性评估指标,采用Elo、HodgeRank和Rank Centrality等聚合方法,超越传统二元成败判断,全面量化生成内容的危害程度。实验表明,该方法持续突破主流LLMs的安全防线,生成高风险内容,为理解越狱攻击风险提供洞见,并助力构建更强防御策略。
原文摘要 · Abstract (English)
Despite the remarkable performance of Large Language Models (LLMs), they remain vulnerable to jailbreak attacks, which can compromise their safety mechanisms. Existing studies often rely on brute-force optimization or manual design, failing to uncover potential risks in real-world scenarios. To address this, we propose a novel jailbreak attack framework, ICRT, inspired by heuristics and biases in human cognition. Leveraging the simplicity effect, we employ cognitive decomposition to reduce the complexity of malicious prompts. Simultaneously, relevance bias is utilized to reorganize prompts, enhancing semantic alignment and inducing harmful outputs effectively. Furthermore, we introduce a ranking-based harmfulness evaluation metric that surpasses the traditional binary success-or-failure paradigm by employing ranking aggregation methods such as Elo, HodgeRank, and Rank Centrality to comprehensively quantify the harmfulness of generated content. Experimental results show that our approach consistently bypasses mainstream LLMs' safety mechanisms and generates high-risk content, providing insights into jailbreak attack risks and contributing to stronger defense strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。