arXiv:2510.12463cs.CL2025-10被引 4

模型语法泛化能力主要受语料丰富度影响,而非语法复杂度。

Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test

  • 用扩展版Wug测试评估六种模型在四种语言上的形态泛化能力
  • 西班牙语和英语准确率显著高于加泰罗尼亚语和希腊语
  • 模型表现更依赖语料规模,非语法结构难度

大型语言模型的语言能力仍存争议。本研究通过多语言版Wug测试,考察模型在新词形态泛化任务中的表现,对比六种模型在加泰罗尼亚语、英语、希腊语和西班牙语四种部分无关语言上的表现,并与人类说话者比较。结果表明,模型能以接近人类的准确率对未见词汇进行形态推断,但其准确率模式更符合语言社区规模和数据可得性,而非语法结构复杂度。拥有更大使用者群体和更强数字表征的西班牙语和英语表现更好,而加泰罗尼亚语和希腊语则相对较低。研究说明模型行为主要受语言资源丰度驱动,其看似类人表现实则表面相似。

原文摘要 · Abstract (English)

The linguistic abilities of Large Language Models are a matter of ongoing debate. This study contributes to this discussion by investigating model performance in a morphological generalization task that involves novel words. Using a multilingual adaptation of the Wug Test, six models were tested across four partially unrelated languages (Catalan, English, Greek, and Spanish) and compared with human speakers. The aim is to determine whether model accuracy approximates human competence and whether it is shaped primarily by linguistic complexity or by the size of the linguistic community, which affects the quantity of available training data. Consistent with previous research, the results show that the models are able to generalize morphological processes to unseen words with human-like accuracy. However, accuracy patterns align more closely with community size and data availability than with structural complexity, refining earlier claims in the literature. In particular, languages with larger speaker communities and stronger digital representation, such as Spanish and English, revealed higher accuracy than less-resourced ones like Catalan and Greek. Overall, our findings suggest that model behavior is mainly driven by the richness of linguistic resources rather than by sensitivity to grammatical complexity, reflecting a form of performance that resembles human linguistic competence only superficially.

语言模型形态学Wug测试资源影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。