揭示语言模型语义不变性的几何结构,可用于零样本模型归属
Invariant Features in Language Models: Geometric Characterization and Model Attribution

- 从局部几何视角分析隐空间,区分语义变化与不变方向
- 发现特定深度区域存在语义不变结构,且其影响因果性显著
- 基于不变特征实现跨模型零样本归属,适用于模型可解释性研究
语言模型对改写具有强鲁棒性,暗示语义信息可能通过稳定内部表征编码,但此类不变性的结构与起源仍不清晰。我们提出一种局部几何框架:语义等价输入在隐空间中占据有结构的区域,改写变化沿干扰方向分布,语义一致性则保留在不变子空间中。在此基础上,本文做出三项贡献:(1) 对不变隐特征的几何刻画;(2) 一种对比子空间发现方法,用于分离语义变化与语义保持的变异;(3) 将不变表示应用于零样本模型归属。在不同模型和层上的实证结果支持上述结论:不变结构出现在特定深度区域,语义位移主要位于干扰子空间之外,表示层面干预表明不变成分对模型输出具有因果作用。不变表示还捕捉到模型特异的几何模式,实现高精度归属。这些发现表明,语义不变性可被视为隐表示的局部几何属性,为语言模型组织意义提供了原理性视角。
原文摘要 · Abstract (English)
Language models exhibit strong robustness to paraphrasing, suggesting that semantic information may be encoded through stable internal representations, yet the structure and origin of such invariance remain unclear. We propose a local geometric framework in which semantically equivalent inputs occupy structured regions in latent space, with paraphrastic variation along nuisance directions and semantic identity preserved in invariant subspaces. Building on this view, we make three contributions: (1) a geometric characterization of invariant latent features, (2) a contrastive subspace discovery method that separates semantic-changing from semantic-preserving variation, and (3) an application of invariant representations to zero-shot model attribution. Across models and layers, empirical results support these contributions. Invariant structure emerges in specific depth regions, semantic displacement lies largely outside the nuisance subspace, and representation-level interventions indicate a causal role of invariant components in model outputs. Invariant representations also capture model-specific geometric patterns, enabling accurate attribution. These findings suggest that semantic invariance can be viewed as a local geometric property of latent representations, offering a principled perspective on how language models organize meaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。