语言模型更难学会违背语言普遍规律的语法,说明学习存在通用偏好。
Can Language Models Learn Typologically Implausible Languages?
- 用自然化的反事实语言测试模型学习能力
- 模型学反向语法慢但最终表现相近
- 支持语言规律源于通用学习偏见
人类语言的语法特征存在有趣的关联,常归因于人类的学习偏好。然而现有证据多来自简化的人工语言实验,难以判断这些关联是源于通用还是语言特异性偏见。语言模型为大规模、高自然度的人工语言学习研究提供了可能。本文首先探讨语言模型如何更好揭示通用学习偏见在语言普遍性中的作用,随后评估语言模型在典型合理与不合理语言(遵循语序普遍规律)间的可学性差异。我们开展对称跨语言实验,训练和测试模型在高度自然化但违背事实的英语(头前置)与日语(头后置)变体上。相比以往工作,我们的数据集更自然,接近可接受性边界。实验表明,模型学习这些细微不合理语言时更慢,但在部分指标上最终表现仍相似。结果支持语言模型表现出一定的类型学对齐学习偏好,提示语言类型规律至少部分源自通用学习偏见。
原文摘要 · Abstract (English)
Grammatical features across human languages show intriguing correlations often attributed to learning biases in humans. However, empirical evidence has been limited to experiments with highly simplified artificial languages, and whether these correlations arise from domain-general or language-specific biases remains a matter of debate. Language models (LMs) provide an opportunity to study artificial language learning at a large scale and with a high degree of naturalism. In this paper, we begin with an in-depth discussion of how LMs allow us to better determine the role of domain-general learning biases in language universals. We then assess learnability differences for LMs resulting from typologically plausible and implausible languages closely following the word-order universals identified by linguistic typologists. We conduct a symmetrical cross-lingual study training and testing LMs on an array of highly naturalistic but counterfactual versions of the English (head-initial) and Japanese (head-final) languages. Compared to similar work, our datasets are more naturalistic and fall closer to the boundary of plausibility. Our experiments show that these LMs are often slower to learn these subtly implausible languages, while ultimately achieving similar performance on some metrics regardless of typological plausibility. These findings lend credence to the conclusion that LMs do show some typologically-aligned learning preferences, and that the typological patterns may result from, at least to some degree, domain-general learning biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。