arXiv:2505.13500cs.CLcs.AI2025-05被引 5

给大模型加噪声,安全防线被轻松攻破

Noise Injection Systemically Degrades Large Language Model Safety Guardrails

  • 向模型激活值系统注入高斯噪声,测试安全机制鲁棒性
  • 有害输出率最高上升27%,深度安全微调无效
  • 适合关注大模型安全漏洞的研究者和部署方

大型语言模型(LLMs)的安全防护机制对防止有害输出至关重要,但其在扰动下的稳定性仍不清楚。本文通过系统性地向多个开源大模型的激活值注入高斯噪声,研究安全微调的鲁棒性。结果表明:(1)高斯噪声使有害输出率提升最高达27%(p < 0.001);(2)更深的安全微调并未提供额外保护;(3)链式思维推理能力基本保持不变。这些发现揭示了当前安全对齐技术的关键脆弱性,提示基于推理和强化学习的方法可能是构建更鲁棒人工智能安全系统的潜在方向。该结果对安全关键场景下大模型的实际部署具有重要启示,表明即使无对抗提示,广泛采用的安全微调方法也可能失效。

原文摘要 · Abstract (English)

Safety guardrails in large language models (LLMs) are a critical component in preventing harmful outputs. Yet, their resilience under perturbation remains poorly understood. In this paper, we investigate the robustness of safety fine-tuning in LLMs by systematically injecting Gaussian noise into model activations. We show across multiple open-weight models that (1) Gaussian noise raises harmful-output rates (p < 0.001) by up to 27%, (2) that deeper safety fine-tuning affords no extra protection, and (3) that chain-of-thought reasoning remains largely intact. The findings reveal critical vulnerabilities in current safety alignment techniques and highlight the potential of reasoning-based and reinforcement learning approaches as promising direction for developing more robust AI safety systems. These results have important implications for real-world deployment of LLMs in safety-critical applications as these results imply that widely-deployed safety tuning methods can fail even without adversarial prompts.

大模型安全噪声攻击安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。