arXiv:2507.14640cs.CL2025-07NAACL被引 1

发现语言模型中形态关系可通过线性变换准确解码,解释力达90%。

Linear Relational Decoding of Morphology in Language Models

  • 用线性映射解码主语与宾语间的形态关系,基于中间层表示与模型梯度。
  • 在多种语言和模型上实现90%的解码忠实度,表现稳定。
  • 揭示了语言模型中概念关系以稀疏线性方式编码,适合可解释性研究者。

研究发现,对某些主宾关系,变压器模型的计算可用两部分仿射近似良好描述。通过改进的更大类比测试集,我们证明:利用主语标记中间层表示 s 和由模型导数推导出的变换矩阵 W 构成的线性变换 Ws,可准确还原多数关系下的最终宾语状态。该线性方法在形态关系上达到90%的忠实度,并在多语言及不同模型中均验证有效。结果表明,语言模型中的某些概念关系(如形态)可从隐空间直接解读,且主要通过跨层稀疏线性变换编码。

原文摘要 · Abstract (English)

A two-part affine approximation has been found to be a good approximation for transformer computations over certain subject object relations. Adapting the Bigger Analogy Test Set, we show that the linear transformation Ws, where s is a middle layer representation of a subject token and W is derived from model derivatives, is also able to accurately reproduce final object states for many relations. This linear technique is able to achieve 90% faithfulness on morphological relations, and we show similar findings multi-lingually and across models. Our findings indicate that some conceptual relationships in language models, such as morphology, are readily interpretable from latent space, and are sparsely encoded by cross-layer linear transformations.

可解释性形态学线性解码语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。