arXiv:2605.18537cs.LGcs.AI2026-05

发现大模型中隐藏的时空特征流形,可直接操控生成内容。

Probing for Representation Manifolds in Superposition

论文配图:Probing for Representation Manifolds in Superposition
图 1 · 摘自论文原文
  • 用流形探测器识别超叠加表示中的可线性预测特征空间
  • 在Llama 2-7b中发现时间与空间的可解释特征流形
  • 沿流形调控能改变模型对作品发布年份的生成结果

本文提出流形探测方法,用于在超叠加表示中发现可解释的特征流形。该方法通过学习概念特征的线性可预测空间及编码方向,扩展了传统线性回归探测器。我们在Llama 2-7b模型中验证了该方法,成功发现了时间与空间的特征流形,其对应特征具有可解释性。进一步实验表明,沿时间流形进行操控,可影响模型对著名歌曲、电影和书籍发布年份的生成结果,证明该流形与模型行为存在因果关联。

原文摘要 · Abstract (English)

This paper introduces the Manifold Probe, a supervised method for discovering representation manifolds in superposition. The method generalizes linear regression probes by learning the space of features of a concept that can be linearly predicted from the representations, and then learning the directions used to encode them. We demonstrate the probe on representations of time and space in Llama 2-7b, finding manifolds which linearly represent an interpretable set of features in each case. In the case of time, we show that by steering along the manifold, we can influence the model's completions about the years in which famous songs, movies and books were released, providing evidence that the Manifold Probe can discover manifolds which are causally involved in model behaviour.

表示学习模型可解释性大模型探针

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。