发现语言模型在预训练中存在关键相变点,导致人类阅读预测能力下降。
Language Models Grow Less Humanlike beyond Phase Transition
- 通过注意力头的快速涌现识别出预训练相变点
- 相变后继续训练会持续损害模型对人类阅读行为的预测能力
- 适用于研究模型对齐与认知模拟的学者
语言模型(LMs)在预训练过程中对人类阅读行为的拟合能力(即心理测量预测力,PPP)通常先提升至一个临界点,之后趋于平稳或下降。尽管已有研究提出词频、注意力中的近期偏差和上下文长度等因素可能影响PPP,但尚无统一解释说明该临界点为何出现及其与预训练动态的关联。本文假设其根本原因是预训练过程中的相变现象,表现为专门化注意力头的快速涌现。我们通过一系列相关性与因果实验验证,该相变正是导致PPP临界点的原因。进一步发现,相变并非直接造成PPP下降,而是改变了模型后续的学习动态,使得继续训练反而持续损害PPP。
原文摘要 · Abstract (English)
LMs' alignment with human reading behavior (i.e. psychometric predictive power; PPP) is known to improve during pretraining up to a tipping point, beyond which it either plateaus or degrades. Various factors, such as word frequency, recency bias in attention, and context size, have been theorized to affect PPP, yet there is no current account that explains why such a tipping point exists, and how it interacts with LMs' pretraining dynamics more generally. We hypothesize that the underlying factor is a pretraining phase transition, characterized by the rapid emergence of specialized attention heads. We conduct a series of correlational and causal experiments to show that such a phase transition is responsible for the tipping point in PPP. We then show that, rather than producing attention patterns that contribute to the degradation in PPP, phase transitions alter the subsequent learning dynamics of the model, such that further training keeps damaging PPP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。