语言模型难以从间接数据中习得语法规则,远不如人类高效。
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
- 通过插入新词的合成数据测试模型对语法的推断能力。
- 即使多次接触结构相同但词汇不同的例子,模型仍无法掌握规则。
- 适合研究语言习得机制与数据效率的学者参考。
语言模型需要多少数据才能推断出句子可接受性?当前模型在数据效率上仍远不及人类。本文探讨语言模型是否能有效利用间接证据(即非直接呈现的语法信息)。人类能高效利用此类证据,这被认为是语言习得高效性的关键归纳偏置。为此,我们提出Wug InDirect Evidence Test(WIDET)数据集,将合成的含新词(wug词)的训练样本注入预训练数据,并在评估阶段考察模型对这些新词句法可接受性的判断。通过控制间接程度和样本数量,实验发现:即使反复接触结构一致但词汇不同的训练实例,模型依然无法推导出语法规则。这一结果提示未来应探索基于潜在间接证据的语法知识诱导方法。
原文摘要 · Abstract (English)
What kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability? Recent language models still have much room for improvement in their data efficiency compared to humans. This paper investigates whether language models efficiently use indirect data (indirect evidence), from which they infer sentence acceptability. In contrast, humans use indirect evidence efficiently, which is considered one of the inductive biases contributing to efficient language acquisition. To explore this question, we introduce the Wug InDirect Evidence Test (WIDET), a dataset consisting of training instances inserted into the pre-training data and evaluation instances. We inject synthetic instances with newly coined wug words into pretraining data and explore the model's behavior on evaluation data that assesses grammatical acceptability regarding those words. We prepare the injected instances by varying their levels of indirectness and quantity. Our experiments surprisingly show that language models do not induce grammatical knowledge even after repeated exposure to instances with the same structure but differing only in lexical items from evaluation instances in certain language phenomena. Our findings suggest a potential direction for future research: developing models that use latent indirect evidence to induce grammatical knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。