Transformer模型学习语言时,抽象能力先于具体词汇出现,揭示了抽象在语言习得中的核心作用。
Humans and transformer LMs: Abstraction drives language learning
- 通过追踪下一个词分布的差异,分析模型训练中抽象与具体行为的演变顺序。
- 在GPT-2 small上发现,类别级抽象行为比具体词项行为更早显现,且不同语言行为分阶段突现。
- 结果支持抽象是语言学习关键机制的观点,对理解人类语言习得有启发意义。
分类是人类语言能力的核心。本文通过比较基于Transformer的语言模型在训练过程中的行为,与人类语言习得中抽象特征驱动和具体实例驱动两种理论,研究其如何形成词义和句法类别。采用新颖的基于分歧的度量方法,追踪下一词分布的学习轨迹。在GPT-2 small的实验中发现:(i) 构造被掌握时,类别级抽象行为出现在词项特定行为之前;(ii) 不同语言行为在训练中分阶段、突现式地出现。结果表明,抽象在语言模型学习中起关键作用,为语言习得模型提供了存在性证据。
原文摘要 · Abstract (English)
Categorization is a core component of human linguistic competence. We investigate how a transformer-based language model (LM) learns linguistic categories by comparing its behaviour over the course of training to behaviours which characterize abstract feature-based and concrete exemplar-based accounts of human language acquisition. We investigate how lexical semantic and syntactic categories emerge using novel divergence-based metrics that track learning trajectories using next-token distributions. In experiments with GPT-2 small, we find that (i) when a construction is learned, abstract class-level behaviour is evident at earlier steps than lexical item-specific behaviour, and (ii) that different linguistic behaviours emerge abruptly in sequence at different points in training, revealing that abstraction plays a key role in how LMs learn. This result informs the models of human language acquisition that LMs may serve as an existence proof for.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。