测试大模型在构词组合上的泛化能力,发现其表现远不如人类。
Evaluating Morphological Compositional Generalization in Large Language Models
- 以词素为基本单元,设计生成与判别任务评估构词能力。
- 面对新词根时模型性能随复杂度上升急剧下降,准确率远低于人类。
- 适合关注语言模型本质理解与人类语言创造力对比的研究者。
大型语言模型(LLMs)在自然语言生成与理解任务中取得了显著进展,但其语言泛化能力仍存疑,难以确定这些模型是否像人类一样学习语言。人类在语言使用中表现出组合性泛化和语言创造力,而大模型在形态学方面是否具备类似能力尚不明确。本文从组合性视角系统研究了大模型的形态学泛化能力,将词素定义为组合基元,并设计了一套新颖的生成与判别任务,以评估其形态生产力与系统性。聚焦于土耳其语、芬兰语等黏着语,我们评估了包括GPT-4和Gemini在内的多个先进指令微调多语言模型。结果表明,当应用于新词根时,大模型在形态组合泛化上表现不佳,随着形态复杂度增加,性能显著下降。尽管模型在识别单个形态组合方面优于随机水平,但缺乏系统性,与人类相比存在显著准确率差距。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks. However, their linguistic generalization capabilities remain questionable, raising doubts about whether these models learn language similarly to humans. While humans exhibit compositional generalization and linguistic creativity in language use, the extent to which LLMs replicate these abilities, particularly in morphology, is under-explored. In this work, we systematically investigate the morphological generalization abilities of LLMs through the lens of compositionality. We define morphemes as compositional primitives and design a novel suite of generative and discriminative tasks to assess morphological productivity and systematicity. Focusing on agglutinative languages such as Turkish and Finnish, we evaluate several state-of-the-art instruction-finetuned multilingual models, including GPT-4 and Gemini. Our analysis shows that LLMs struggle with morphological compositional generalization particularly when applied to novel word roots, with performance declining sharply as morphological complexity increases. While models can identify individual morphological combinations better than chance, their performance lacks systematicity, leading to significant accuracy gaps compared to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。