arXiv:2511.16467cs.CLcs.AI2025-11被引 1

揭示大模型处理成语时的隐藏计算机制。

Anatomy of an Idiom: Tracing Non-Compositionality in Language Models

  • 用路径修补法发现模型内部处理成语的专用计算电路。
  • 识别出频繁激活的'成语头'和增强的词语注意力模式。
  • 为理解语言非组合性提供新视角,适合关注模型推理机制的研究者。

我们采用一种新型电路发现与分析技术,研究基于Transformer的语言模型对习语表达的处理方式。通过改进的路径修补算法首次发现处理习语时具有独特的计算模式。识别出在多种习语中频繁激活的注意力头(称为'成语头'),以及因早期处理导致的习语词元间增强注意力,称之为'增强接收'。分析这些现象及其电路特征,揭示了变压器如何在计算效率与鲁棒性之间取得平衡。研究成果为理解变压器处理非组合性语言提供了洞见,并为解析更复杂的语法结构提供了可行路径。

原文摘要 · Abstract (English)

We investigate the processing of idiomatic expressions in transformer-based language models using a novel set of techniques for circuit discovery and analysis. First discovering circuits via a modified path patching algorithm, we find that idiom processing exhibits distinct computational patterns. We identify and investigate ``Idiom Heads,'' attention heads that frequently activate across different idioms, as well as enhanced attention between idiom tokens due to earlier processing, which we term ``augmented reception.'' We analyze these phenomena and the general features of the discovered circuits as mechanisms by which transformers balance computational efficiency and robustness. Finally, these findings provide insights into how transformers handle non-compositional language and suggest pathways for understanding the processing of more complex grammatical constructions.

语言模型习语理解注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。