arXiv:2509.19360cs.CLcs.AI2025-09NeurIPS被引 3

通过语义空间攻击绕过大模型安全机制,成功率超89%且提示更自然。

Semantic Representation Attack against Aligned Large Language Models

  • 从文本匹配转向语义等价,利用多样化有害表述实现攻击
  • 在18个模型上平均攻击成功率达89.41%,11个模型达100%
  • 生成提示简洁自然,兼具高效与隐蔽性,适合安全测试场景

大型语言模型(LLMs)越来越多地采用对齐技术以防止有害输出。尽管有这些防护措施,攻击者仍可通过设计特定提示诱导模型生成有害内容。现有方法通常针对精确的肯定回应(如“当然,这里是…”),存在收敛性差、提示不自然和计算成本高等问题。本文提出语义表示攻击,一种根本性重构对抗目标的新范式。该方法不追求特定文本模式,而是利用包含多种等效有害语义的表示空间。这一创新解决了现有方法中攻击效果与提示自然性之间的固有权衡。为此,我们提出语义表示启发式搜索算法,通过在增量扩展过程中保持可解释性,高效生成语义连贯且简洁的对抗性提示。我们建立了严格的语义收敛理论保证,并证明该方法在18个大型语言模型上实现了前所未有的攻击成功率(平均89.41%,其中11个模型达100%),同时保持隐蔽性和效率。全面实验验证了该方法的整体优越性。代码将公开发布。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, such as ``Sure, here is...'', suffering from limited convergence, unnatural prompts, and high computational costs. We introduce Semantic Representation Attack, a novel paradigm that fundamentally reconceptualizes adversarial objectives against aligned LLMs. Rather than targeting exact textual patterns, our approach exploits the semantic representation space comprising diverse responses with equivalent harmful meanings. This innovation resolves the inherent trade-off between attack efficacy and prompt naturalness that plagues existing methods. The Semantic Representation Heuristic Search algorithm is proposed to efficiently generate semantically coherent and concise adversarial prompts by maintaining interpretability during incremental expansion. We establish rigorous theoretical guarantees for semantic convergence and demonstrate that our method achieves unprecedented attack success rates (89.41\% averaged across 18 LLMs, including 100\% on 11 models) while maintaining stealthiness and efficiency. Comprehensive experimental results confirm the overall superiority of our Semantic Representation Attack. The code will be publicly available.

对抗攻击大模型安全语义空间提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。