发现大模型生成词汇靠类比而非规则,尤其在不规则词上更明显。
Derivational Morphology Reveals Analogical Generalization in Large Language Models
- 用类比和规则模型对比大模型对新词的生成,分析其内在机制。
- 对不规则名词化形式,类比模型预测更接近大模型表现。
- 高频词影响模型行为,支持类比学习而非规则学习的假设。
大语言模型的语言泛化机制是什么?现有研究多关注规则性语言现象,但无法区分规则与类比机制的差异。本文聚焦英语形容词名词化这一具显著变异性的问题,以GPT-J为例,通过拟合规则与类比认知模型,对比其在伪词上的预测能力。结果表明:对于规则性形式,两类模型解释力相当;但对于变异性形式,类比模型显著更优。此外,即使对规则形式,大模型对词频敏感,这符合类比机制特征,而规则模型无法解释。研究驳斥了大模型依赖规则的假设,表明其语言泛化主要基于存储实例间的相似性操作。
原文摘要 · Abstract (English)
What mechanisms underlie linguistic generalization in large language models (LLMs)? This question has attracted considerable attention, with most studies analyzing the extent to which the language skills of LLMs resemble rules. As of yet, it is not known whether linguistic generalization in LLMs could equally well be explained as the result of analogical processes, which can be formalized as similarity operations on stored exemplars. A key shortcoming of prior research is its focus on linguistic phenomena with a high degree of regularity, for which rule-based and analogical approaches make the same predictions. Here, we instead examine derivational morphology, specifically English adjective nominalization, which displays notable variability. We introduce a new method for investigating linguistic generalization in LLMs: focusing on GPT-J, we fit cognitive models that instantiate rule-based and analogical learning to the LLM training data and compare their predictions on a set of nonce adjectives with those of the LLM, allowing us to draw direct conclusions regarding underlying mechanisms. As expected, rule-based and analogical models explain the predictions of GPT-J equally well for adjectives with regular nominalization patterns. However, for adjectives with variable nominalization patterns, the analogical model provides a much better match. Furthermore, GPT-J's behavior is sensitive to the individual word frequencies, even for regular forms, a behavior that is consistent with an analogical account of regular forms but not a rule-based one. These findings refute the hypothesis that GPT-J's linguistic generalization on adjective nominalization involves rules, suggesting similarity operations on stored exemplars as the underlying mechanism. Overall, our study suggests that analogical processes play a bigger role in the linguistic generalization of LLMs than previously thought.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。