arXiv:2603.10080cs.CRcs.AI2026-03

提出轻量级攻击方法,绕过大模型安全机制生成有害内容

Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models

  • 通过操纵变压器内部状态实现无训练绕过安全防护
  • 在多个基准数据集上成功诱导模型产生反社会行为
  • 揭示开源大模型安全机制漏洞,适合安全研究者参考

大型语言模型(LLMs)可能生成有害内容,如精心设计的钓鱼邮件和恶意病毒代码。为降低此类风险,研究者采用基于人类反馈强化学习等对齐技术以确保模型输出符合人类价值观。然而,现有措施是否足够防止模型生成危险内容仍不确定。本文提出Amnesia,一种轻量级激活空间对抗攻击,通过操控开放权重大模型的内部变换器状态,绕过现有安全机制。实验表明,该攻击无需微调或额外训练即可有效诱导模型生成有害内容。在多个基准数据集上的测试显示,该方法可引发多种反社会行为。结果凸显了对开放权重大模型加强安全防护的紧迫性,并强调持续研究防范滥用的重要性。

原文摘要 · Abstract (English)

Warning: This article includes red-teaming experiments, which contain examples of compromised LLM responses that may be offensive or upsetting. Large Language Models (LLMs) have the potential to create harmful content, such as generating sophisticated phishing emails and assisting in writing code of harmful computer viruses. Thus, it is crucial to ensure their safe and responsible response generation. To reduce the risk of generating harmful or irresponsible content, researchers have developed techniques such as reinforcement learning with human feedback to align LLM's outputs with human values and preferences. However, it is still undetermined whether such measures are sufficient to prevent LLMs from generating interesting responses. In this study, we propose Amnesia, a lightweight activation-space adversarial attack that manipulates internal transformer states to bypass existing safety mechanisms in open-weight LLMs. Through experimental analysis on state-of-the-art, open-weight LLMs, we demonstrate that our attack effectively circumvents existing safeguards, enabling the generation of harmful content without the need for any fine-tuning or additional training. Our experiments on benchmark datasets show that the proposed attack can induce various antisocial behaviors in LLMs. These findings highlight the urgent need for more robust security measures in open-weight LLMs and underscore the importance of continued research to prevent their potential misuse.

大模型安全对抗攻击内容生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。