用思维链引导模型预判攻击,实现安全与自然对话的平衡
Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
- 通过思维链自生成对抗场景并制定防御策略
- 近零越狱成功率,毒性极低,拒绝率低于4%
- 适合追求高安全与流畅交互的AI应用开发者
大型语言模型虽能力出色,但仍面临越狱、有毒内容、幻觉和偏见等多重安全风险。现有防御手段多仅针对单一威胁或采取僵硬拒绝策略,牺牲用户体验且难以泛化至新型攻击。本文提出对抗场景外推(ASE)框架,利用思维链(CoT)推理,在生成回复前引导模型自主构思潜在攻击场景并制定防御方案。在四个对抗基准上对四款最新LLM的全面评估显示,ASE实现近乎零的越狱成功率和极低毒性,同时将直接拒绝率降至4%以下;在鲁棒性与自然性权衡上优于六种前沿防御方法,对抗问答准确率达92%-99%,偏见评分降低4-10倍。该方法将对抗感知转化为内在认知过程,为安全而自然的人机交互树立新范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user experience and failing to generalize across diverse and novel attacks. This paper introduces Adversarial Scenario Extrapolation (ASE), a novel inference-time computation framework that leverages Chain-of-Thought (CoT) reasoning to simultaneously enhance LLM robustness and seamlessness. ASE guides the LLM through a self-generative process of contemplating potential adversarial scenarios and formulating defensive strategies before generating a response to the user query. Comprehensive evaluation on four adversarial benchmarks with four latest LLMs shows that ASE achieves near-zero jailbreak attack success rates and minimal toxicity, while slashing outright rejections to <4%. ASE outperforms six state-of-the-art defenses in robustness-seamlessness trade-offs, with 92-99% accuracy on adversarial Q&A and 4-10x lower bias scores. By transforming adversarial perception into an intrinsic cognitive process, ASE sets a new paradigm for secure and natural human-AI interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。