语言模型在训练早期突然学会原文检索,且与复杂任务能力正相关。
Transformer verbatim in-context retrieval across time and scale
- 训练初期(约1%令牌后)出现突然的原文检索能力跃升。
- 检索能力与零样本基准表现正相关,且对具体名词更优。
- 适合关注模型内在机制与认知类比的研究者阅读。
为预测后续文本,语言模型有时需原样检索上下文信息。本报告研究了语言模型在训练过程中(跨时间)及规模增大时(跨尺度)检索任意上下文名词的能力发展。我们进一步探讨了该能力是否与更复杂的零样本基准学习相关。受人类短期记忆语义效应启发,评估了目标名词的语义特征——即其指代具体或抽象实体的程度(由人类标注)。结果表明,原文检索能力在训练早期(约1%训练令牌后)突然形成,此现象在从14M至12B参数的各模型中均存在,最小两模型略晚发生。所有模型在跃迁点附近均表现出对具体名词优于抽象名词的检索优势,但除两个最小模型外,该优势在训练后期逐渐消失。
原文摘要 · Abstract (English)
To predict upcoming text, language models must in some cases retrieve in-context information verbatim. In this report, we investigated how the ability of language models to retrieve arbitrary in-context nouns developed during training (across time) and as language models trained on the same dataset increase in size (across scale). We then asked whether learning of in-context retrieval correlates with learning of more challenging zero-shot benchmarks. Furthermore, inspired by semantic effects in human short-term memory, we evaluated the retrieval with respect to a major semantic component of target nouns, namely whether they denote a concrete or abstract entity, as rated by humans. We show that verbatim in-context retrieval developed in a sudden transition early in the training process, after about 1% of the training tokens. This was observed across model sizes (from 14M and up to 12B parameters), and the transition occurred slightly later for the two smallest models. We further found that the development of verbatim in-context retrieval is positively correlated with the learning of zero-shot benchmarks. Around the transition point, all models showed the advantage of retrieving concrete nouns as opposed to abstract nouns. In all but two smallest models, the advantage dissipated away toward the end of training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。