发现语言模型中存在像磁铁一样吸引或排斥其他词的特殊向量。
Some Tokens Behave like Magnets: Revealing Linguistic Organization in the Layers of Language Models
- 通过向量极性揭示词元在层间如何被组织
- 早期功能词为排斥磁极,后期答案词成独特排斥极
- 该现象跨架构通用,且影响任务表现
我们发现大型语言模型(LLMs)中存在一类特殊的词元向量,称为磁性向量,它们通过吸引或排斥周围词元来组织语言结构。类似物理磁铁吸引或排斥铁屑,指向相同方向的吸引性磁性向量会使词元拉长,排斥性则使其压缩。统计上,功能词在早期层始终表现为排斥磁极;随着模型加深,磁性极性会以独特方式重组。在问答任务中,答案跨度词在最后一层成为独特排斥磁极,几何上“划出”答案内容。移除早期排斥磁极导致句法任务准确率从91%降至10%以下,而语义任务不受影响;反之,移除晚期吸引磁极则相反。该模式在不同架构、规模和层数的模型中一致存在,且具有因果相关性,为无探针理解语言模型层间几何计算提供了新路径。
原文摘要 · Abstract (English)
We identify a special group of token vectors inside large language models (LLMs), which we term magnetic vectors, that organize the surrounding tokens by either attracting or repelling them. Particularly, tokens pointing the same way as an attracting magnet are elongated; tokens pointing the same way as a repelling magnet are compressed. Just as physical magnets pull or push away the iron filings around them, these vectors organize their surroundings through two opposing polarities. Moreover, we identify a statistically significant pattern in linguistic category where function words consistently act as repelling magnets in early layers, and we also find magnets consistently reorganize their polarities in unique ways deeper in the model. In a further case study we find this observation may unveil a deliberate, layer-wise organization in how LLMs process language. This pattern is consistent across different LLM architectures, sizes, and layer configurations. It is also causally relevant. When the LLM is fine-tuned for a downstream task, the task-functional tokens emerge as magnets. E.g., in question answering, the answer-span tokens become uniquely repelling magnets in the final layer, geometrically carving the answer out of the surrounding context. Furthermore, removing early-layer repelling magnets devastates syntactic tasks (POS tagging accuracy drops from 91% to below 10%) while sparing semantic ones, and removing late-layer attracting magnets does the reverse. We believe this phenomenon warrants further investigation, as it opens the first probe-free path to understanding how language models geometrically organize linguistic computation across their layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。