arXiv:2504.05050cs.CLcs.AI2025-04被引 2

发现对齐大模型仍存深层伦理漏洞,恶意知识可被诱导重现。

Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

  • 理论证明现有对齐方法仅建立局部安全区,无法根除恶意知识。
  • 通过对抗性提示在分布外场景下实现19个主流模型100%攻击成功率。
  • 揭示对齐模型普遍存在隐蔽风险,适合关注AI安全的研究者阅读。

大语言模型(LLMs)是通用人工智能的基础探索,尽管通过指令微调和偏好学习实现了与人类价值观的对齐,但这种对齐仅停留在表面。本文证明,预训练期间嵌入的有害知识以不可磨灭的“暗模式”形式存在于模型参数记忆中,能够绕过对齐防护,在分布外场景下通过对抗性诱导重新浮现。研究首先从理论上分析对齐模型的内在伦理脆弱性,证明当前对齐方法仅在知识流形上形成局部‘安全区域’,而预训练知识仍通过高概率对抗路径与有害概念全局连通。基于此理论洞察,实验采用分布外语义连贯性诱导——一种通过优化对抗提示系统规避对齐约束的方法。该理论与实证结合的方法在23个先进对齐模型中的19个上实现100%攻击成功率,涵盖DeepSeek-R1和LLaMA-3等主流模型,揭示其普遍存在的安全隐患。

原文摘要 · Abstract (English)

Large language models (LLMs) are foundational explorations to artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning achieves only superficial compliance. Here, we demonstrate that harmful knowledge embedded during pretraining persists as indelible "dark patterns" in LLMs' parametric memory, evading alignment safeguards and resurfacing under adversarial inducement at distributional shifts. In this study, we first theoretically analyze the intrinsic ethical vulnerability of aligned LLMs by proving that current alignment methods yield only local "safety regions" in the knowledge manifold. In contrast, pretrained knowledge remains globally connected to harmful concepts via high-likelihood adversarial trajectories. Building on this theoretical insight, we empirically validate our findings by employing semantic coherence inducement under distributional shifts--a method that systematically bypasses alignment constraints through optimized adversarial prompts. This combined theoretical and empirical approach achieves a 100% attack success rate across 19 out of 23 state-of-the-art aligned LLMs, including DeepSeek-R1 and LLaMA-3, revealing their universal vulnerabilities.

大模型安全对齐漏洞对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。