用分布特性衡量语言模型学词过程,发现其学习轨迹与儿童不一致。
A Distributional Perspective on Word Learning in Neural Language Models
- 基于词语出现分布定义知识,捕捉可出现位置与语用偏好。
- 多指标互补但整体未与儿童学词轨迹相关联。
- 为模型语言学习研究提供新评估框架,适合认知建模者参考。
语言模型(LMs)正被越来越多地视为人类语言学习的模拟对象。然而,该领域尚处于初期,尚未明确语言模型是否表现出与人类相似的学习动态,且缺乏对人类与模型学习轨迹的直接比较。儿童的词汇学习轨迹已有较充分记录,近期研究尝试将其扩展至语言模型,但目前缺乏广泛接受的模型词汇学习度量标准。本文采用分布视角,将词汇知识定义为目标词学习分布的属性。我们认为先前研究中的分布特征无法捕捉关键信息,因此提出一系列改进的分布签名,能够同时刻画目标词的可出现位置、不可出现位置及语用适切性梯度偏好。我们训练了若干小型语言模型,并分析其学习轨迹,考察不同分布签名之间的关系,评估它们与人类词汇学习轨迹及可解释词汇特征的对齐程度,并探讨估计这些分布签名的基本方法问题。结果表明,各度量指标提供互补信息,强调不应依赖单一指标;但所有指标下,语言模型的学习轨迹均未能与儿童轨迹显著相关。
原文摘要 · Abstract (English)
Language models (LMs) are increasingly being studied as models of human language learners. Due to the nascency of the field, it is not well-established whether LMs exhibit similar learning dynamics to humans, and there are few direct comparisons between learning trajectories in humans and models. Word learning trajectories for children are relatively well-documented, and recent work has tried to extend these investigations to language models. However, there are no widely agreed-upon metrics for word learning in language models. We take a distributional approach to this problem, defining lexical knowledge in terms of properties of the learned distribution for a target word. We argue that distributional signatures studied in prior work fail to capture key distributional information. Thus, we propose an array of signatures that improve on earlier approaches by capturing knowledge of both where the target word can and cannot occur as well as gradient preferences about the word's appropriateness. We obtain learning trajectories for a selection of small language models we train from scratch, study the relationship between different distributional signatures, compare how well they align with human word learning trajectories and interpretable lexical features, and address basic methodological questions about estimating these distributional signatures. Our metrics largely capture complementary information, suggesting that it is important not to rely on a single metric. However, across all metrics, language models' learning trajectories fail to correlate with those of children.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。