arXiv:2605.21391cs.CL2026-05

发现解码器模型处理隐喻时出现多尺度协同激活现象。

Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy

论文配图:Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy
图 1 · 摘自论文原文
  • 提出条件尺度熵衡量模型层间跨频率尺度的计算广度。
  • 隐喻词在连续层中谱宽显著高于字面词,覆盖124M至20B参数模型。
  • 该现象与语义复杂度无关,适用于多种主流解码器架构。

隐喻要求语言模型重新解释一个上下文含义偏离其字面意义的词。理解变换器模型如何在深度层级上组织这种再解释,仍是机制可解释性中的开放问题。本文引入条件尺度熵(CSE),一种基于小波的度量方法,用于衡量变换器在每一层位置上的计算在频率尺度上的广泛参与程度。两个定理表明,CSE对更新幅度不变,从而将更新的结构模式与其强度分离。利用CSE,我们发现在从124M到20B参数的所有测试解码器架构(GPT-2系列、LLaMA-2 7B、GPT-oss 20B)中,隐喻词在连续层位置上产生的谱宽显著高于字面词。该效应通过基于聚类的置换校正,且在不同模型的早期到中期相对深度范围内重现,同时与独立分析的200组自然语境下隐喻对(VUA pairs)结果一致。特定性控制显示,该效应并非由语义复杂度或匹配命题内容解释。这些结果揭示了多尺度协调是所考察解码器架构中隐喻语言处理的一致特征,并确立了CSE作为刻画变换器跨深度结构的原理性工具。

原文摘要 · Abstract (English)

Metaphor requires a language model to resolve a token whose contextual meaning diverges from its basic literal sense. Understanding how transformer models organize this reinterpretation across depth remains an open problem in mechanistic interpretability. We introduce conditional scale entropy (CSE), a wavelet-derived measure of how broadly transformer computation engages across frequency scales at each layer position. Two theorems establish that CSE is invariant to update magnitude, isolating the structural pattern of updates from their intensity. Using CSE, we find that metaphorical tokens produce significantly higher spectral breadth than literal tokens at contiguous layer positions on every decoder-only architecture tested, from 124M to 20B parameters (GPT-2 family, LLaMA-2 7B, GPT-oss 20B). The effect survives cluster-based permutation correction, recurs in the early-to-mid relative depth range across models, and converges with an independent analysis of 200 naturalistic VUA pairs. Specificity controls further show that the effect is not explained by semantic complexity or by matched propositional content. These results identify multi-scale coordination as a consistent signature of metaphorical language processing in the decoder-only architectures examined, and establish CSE as a principled tool for characterizing cross-depth structure in transformers.

隐喻理解可解释性变换器机制多尺度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。