注意力机制通过中间层聚集,让隐变量变为可表述状态。
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

- 隐变量通过中层注意力窗口的聚集被激活,而非通过门控机制。
- 在主检查点上,概念可见性提升0.050百分位,所有测试均显著上升。
- 该机制依赖任务需求,适合研究模型内部表示动态的学者。
语言模型能以可报告形式存储隐变量,且当任务需要灵活复用时,该形式更明显。我们通过雅各比透镜分析开放权重模型,在共享相同上下文的五个任务分支上测试,发现无明确门控机制。任务需求使概念的透镜可见性提升0.050([+0.045, +0.057])百分位,所有四个测量点均呈正向,即使某分支已达到准确率天花板,对比组也更强。一个共享线性映射可在所有分支中解码该变量,性能达选择修正基线的6.4–9.0倍。变量在查询位置的可读态由中层注意力窗口内的聚集决定:将补丁深度与读出深度分离后,非饱和读出下该窗口内传输效率至少高出17倍,且测试中无任一MLP输出有正向贡献。在饱和百分位排名下,该窗口无法定位,此为量测方法的固有特性。无需使用该变量的任务分支仅激发其七分之一强度,表明窗口具有需求特异性。该窗口存在两个可测边界:下方失效、上方破坏,且在64层混合模型与62层密集模型中均位于相同比例深度。我们定位了变量的安装与读取位置,而非路径传输,因路径本身不传递信息。但读出并非使用程度的校准度量:三个组件将其拉近至彼此12%以内,却在答案上造成7.4倍差异。
原文摘要 · Abstract (English)
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。