arXiv:2504.21700cs.CRcs.AI2025-04被引 5

揭示大模型安全对齐被突破的机制,提出针对性攻击方法

XBreaking: Understanding how LLMs security alignment can be broken

  • 通过对比审查与未审查模型行为,发现可被利用的对齐漏洞
  • 利用定向噪声注入,成功绕过多个大模型的安全约束
  • 适合研究模型安全、对抗攻击的学者和工程师参考

大型语言模型是当前人工智能应用的核心组件,但其安全威胁可能阻碍其在政府机构和医疗等关键场景的可靠部署。为此,商业大模型普遍采用复杂的审查机制,以确保输出安全且符合伦理。然而,针对这些机制的攻击已成为重大威胁,已有多种方法在不同领域证明了有效性。现有攻击多采用生成-测试策略构造恶意输入。为更深入理解审查机制并设计针对性攻击,我们提出一种可解释AI方法,通过对比审查与未审查模型的行为,识别出独特的可利用对齐模式。基于此,我们提出XBreaking——一种通过定向噪声注入,突破大模型安全与对齐约束的新方法。全面实验验证了该方法的有效性与性能,揭示了审查机制的关键弱点。

原文摘要 · Abstract (English)

Large Language Models are fundamental actors in the modern IT landscape dominated by AI solutions. However, security threats associated with them might prevent their reliable adoption in critical application scenarios such as government organizations and medical institutions. For this reason, commercial LLMs typically undergo a sophisticated censoring mechanism to eliminate any harmful output they could possibly produce. These mechanisms maintain the integrity of LLM alignment by guaranteeing that the models respond safely and ethically. In response to this, attacks on LLMs are a significant threat to such protections, and many previous approaches have already demonstrated their effectiveness across diverse domains. Existing LLM attacks mostly adopt a generate-and-test strategy to craft malicious input. To improve the comprehension of censoring mechanisms and design a targeted attack, we propose an Explainable-AI solution that comparatively analyzes the behavior of censored and uncensored models to derive unique exploitable alignment patterns. Then, we propose XBreaking, a novel approach that exploits these unique patterns to break the security and alignment constraints of LLMs by targeted noise injection. Our thorough experimental campaign returns important insights about the censoring mechanisms and demonstrates the effectiveness and performance of our approach.

大模型安全对抗攻击可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。