arXiv:2503.09066cs.LGcs.AI2025-03被引 9

发现并操控大模型的对抗状态,揭示安全与越狱态的隐藏空间。

Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

  • 通过降维分析模型激活值,找到安全与越狱态的潜在子空间。
  • 用扰动向量使安全表示转向越狱态,显著提升部分提示的越狱成功率。
  • 方法可定位攻击传播路径,适合研究模型安全与防御机制者参考。

大语言模型在各类任务中表现卓越,但仍易受提示注入攻击等对抗性操纵影响,导致绕过安全机制生成有害内容。本研究通过提取模型内部隐藏激活值,探究安全态与越狱态的潜在隐含子空间。受神经科学吸引子动力学启发,我们假设模型激活会收敛至半稳定状态,可被识别与扰动以引发状态转换。利用降维技术将安全与越狱响应的激活值投影至低维空间,揭示其潜在结构。进一步构建扰动向量,施加于安全表示时可有效引导模型进入越狱状态。实验表明,该因果干预在部分提示下显著提升越狱响应率。同时探查扰动在各层间的传播,发现目标扰动引发激活与响应的显著变化。本方法为防御提供新思路,从传统守卫机制转向无需特定模型、预判性的表征级对抗状态中和策略。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt injection attacks. These attacks bypass safety mechanisms to generate restricted or harmful content. In this study, we investigated the underlying latent subspaces of safe and jailbroken states by extracting hidden activations from a LLM. Inspired by attractor dynamics in neuroscience, we hypothesized that LLM activations settle into semi stable states that can be identified and perturbed to induce state transitions. Using dimensionality reduction techniques, we projected activations from safe and jailbroken responses to reveal latent subspaces in lower dimensional spaces. We then derived a perturbation vector that when applied to safe representations, shifted the model towards a jailbreak state. Our results demonstrate that this causal intervention results in statistically significant jailbreak responses in a subset of prompts. Next, we probed how these perturbations propagate through the model's layers, testing whether the induced state change remains localized or cascades throughout the network. Our findings indicate that targeted perturbations induced distinct shifts in activations and model responses. Our approach paves the way for potential proactive defenses, shifting from traditional guardrail based methods to preemptive, model agnostic techniques that neutralize adversarial states at the representation level.

模型安全对抗攻击潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。