arXiv:2605.01844cs.CL2026-05被引 3

提出柱状表征假说,解释大模型控语为何不稳定。

The Cylindrical Representation Hypothesis for Language Model Steering

论文配图:The Cylindrical Representation Hypothesis for Language Model Steering
图 1 · 摘自论文原文
  • 用柱状结构替代传统正交假设,更贴近真实模型表征。
  • 发现控制敏感区仅在特定方向有效,其他方向可能抑制或延迟效果。
  • 解释了即使方向对齐仍波动的原因,适合研究可控生成的读者。

控语是广泛使用的大型语言模型控制技术,但其效果常不稳定且难以预测。现有理论主要基于线性表征假说(LRH),该假说假设概念可正交化实现无损控制,但在真实表征中这一理想映射失效,无法解释控语的不可预测性。通过放宽LRH的正交性假设而保留线性结构,我们发现概念贡献重叠自然形成样本特异的轴正交结构。为此提出柱状表征假说(CRH):中心轴捕捉概念有无的核心差异并驱动概念生成;周围法平面决定控语敏感度,其中仅特定敏感区域能强激活概念,其余区域可能抑制或延迟激活。尽管法平面可从差向量可靠识别,敏感区域却无法确定,导致层级上的内在不确定性。这一不确定性为控语结果在方向对齐时仍波动提供了合理解释。实验验证了柱状结构的存在,并证明CRH在真实场景中可有效解释控语行为:https://github.com/mbzuai-nlp/CRH。

原文摘要 · Abstract (English)

Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical accounts are largely based on the Linear Representation Hypothesis (LRH). While LRH assumes that concepts can be orthogonalized for lossless control, this idealized mapping fails in real representations and cannot account for the observed unpredictability of steering. By relaxing LRH's orthogonality assumption while preserving linear representations, we show that overlapping concept contributions naturally yield a sample-specific axis-orthogonal structure. We formalize this as the Cylindrical Representation Hypothesis (CRH). In CRH, a central axis captures the main difference between concept absence and presence and drives concept generation. A surrounding normal plane controls steering sensitivity by determining how easily the axis can activate the target concept. Within this plane, only specific sensitive sectors strongly facilitate concept activation, while other sectors can suppress or delay it. While the surrounding normal plane can be reliably identified from difference vectors, the sensitive sector cannot, introducing intrinsic uncertainty at the sector level. This uncertainty provides a principled explanation for why steering outcomes often fluctuate even when using well-aligned directions. Our experiments verify the existence of the cylindrical structure and demonstrate that CRH provides a valid and practical way to interpret model steering behavior in real settings: https://github.com/mbzuai-nlp/CRH.

语言模型控语表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。