arXiv:2507.09709cs.CLcs.LG2025-07被引 13

LLM的语义表示在隐空间中呈线性可分,越深层越明显。

Large Language Models Encode Semantics and Alignment in Linearly Separable Representations

  • 通过分析11个自回归模型,发现语义信息集中于低维子空间且线性可分。
  • 深层特征对结构化推理和对齐行为更敏感,即使表面内容不变。
  • 提出轻量级探针可有效识别恶意查询,超越内置安全机制。

理解大语言模型(LLM)隐空间几何结构是解释其行为和提升对齐性的关键。然而,目前尚不明确LLM在多大程度上以线性方式组织与语义理解相关的表征。为此,我们对11个自回归模型在六个科学领域的隐藏表征进行了大规模实证研究。结果表明,高层语义信息始终存在于跨领域线性可分的低维子空间中。这种可分性在深层网络中更为显著,且在激发结构化推理或对齐行为的提示下尤为突出,即使表面内容保持不变。这些发现推动了直接在隐空间中操作的几何感知工具的发展,用于检测和缓解有害及对抗性内容。作为概念验证,我们在最终层隐藏状态上训练了一个MLP探针,用作轻量级隐空间防护机制。该方法显著提升了对恶意查询和绕过模型内置安全及外部标记级过滤器的提示注入攻击的拒绝率。

原文摘要 · Abstract (English)

Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs linearly organize representations related to semantic understanding. To explore this, we conduct a large-scale empirical study of hidden representations in 11 autoregressive models across six scientific topics. We find that high-level semantic information consistently resides in low-dimensional subspaces that form linearly separable representations across domains. This separability becomes more pronounced in deeper layers and under prompts that elicit structured reasoning or alignment behavior$\unicode{x2013}$even when surface content remains unchanged. These findings motivate geometry-aware tools that operate directly in latent space to detect and mitigate harmful and adversarial content. As a proof of concept, we train an MLP probe on final-layer hidden states as a lightweight latent-space guardrail. This approach substantially improves refusal rates on malicious queries and prompt injections that bypass both the model's built-in safety alignment and external token-level filters.

大模型隐空间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。