让模型像人一样用几例快速学会新词
Rapid Word Learning Through Meta In-Context Learning
- 用特殊占位符训练模型自动生成新词用例
- 在儿童语料上训练后,少样本学词能力媲美大模型
- 适合提升模型对新词的识别与生成能力
人类能从少数例句中快速学会一个新词,并灵活应用于新语境。然而当前语言模型在少样本词学习方面的能力及其提升方法仍研究不足。本文提出一种新方法——元上下文词学习(Minnow),通过特殊占位符让模型基于少量上下文例句生成新词的使用范例,并在大量新词上反复训练以发展通用词学习能力。在人类规模的儿童导向语言数据上从头训练的模型,其少样本词学习能力可媲美预训练于海量数据的大语言模型(LLM)。进一步的判别与生成评估表明,用Minnow微调预训练大模型,可显著提升其对新词的辨识、句法类别判断,以及基于一两个例句生成合理用法和定义的能力。这些结果凸显了Minnow的数据效率及其在词学习任务中的潜力。
原文摘要 · Abstract (English)
Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts. Yet the abilities of current language models for few-shot word learning, and methods for improving these abilities, are underexplored. In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow). This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word. This training is repeated on many new words to develop a general word-learning ability. We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pre-trained on orders of magnitude more data. Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples. These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。