测试大模型对全球语言结构的元语言知识,发现其表现受数据资源影响大。
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
- 基于世界语言结构地图构建多语言问答评测集
- GPT-4o准确率仅0.367,多数模型未超越随机水平
- 低资源语言表现差,数字存在度决定模型评估精度
大语言模型常被评估语言使用能力,但其对语言结构的显式知识仍不清楚。现有语言学基准多聚焦少数现象,强调高资源语言,且极少测试元语言知识——即对语言结构的显性推理。本文基于世界语言结构地图(WALS),构建跨语言元语言知识评测,涵盖2,660种语言的192项语言特征。将WALS特征转化为自然语言多选题,评估模型在各语言上的表现。采用准确率与宏F1,对比随机基线和多数类基线,分析不同语言领域与语言相关因素的影响。结果显示,元语言知识有限:尽管GPT-4o表现最佳,准确率仅为0.367,开源模型更差;所有模型虽高于随机,但未超越多数类基线,表明其仅捕捉泛化模式而缺乏细粒度区分。性能随领域变化,部分反映网络可见性差异。在语言层面,准确率与数字语言地位正相关:数字存在度高、资源丰富的语言得分更高,低资源语言表现较差。预测因子分析确认,资源指标(如维基百科规模、语料可用性)比地理、谱系或社会语言学因素更具解释力。总体而言,大模型的元语言知识呈现碎片化,主要受数据可得性驱动,而非普遍语法能力。我们已将该评测集开源,以支持跨语言评估,并呼吁未来模型增强全球语言多样性。
原文摘要 · Abstract (English)
LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structure remains poorly understood. Existing linguistic benchmarks focus on narrow phenomena, emphasize high-resource languages, and rarely test metalinguistic knowledge - explicit reasoning about language structure. We present a multilingual evaluation of metalinguistic knowledge in LLMs, based on the World Atlas of Language Structures (WALS), documenting 192 linguistic features across 2,660 languages. We convert WALS features into natural-language multiple-choice questions and evaluate models across documented languages. Using accuracy and macro F1, and comparing to chance and majority-class baselines, we assess performance and analyse variation across linguistic domains and language-related factors. Results show limited metalinguistic knowledge: GPT-4o performs best but achieves moderate accuracy (0.367), while open-source models lag. Although all models perform above chance, they fail to outperform the majority-class baseline, suggesting they capture broad cross-linguistic patterns but lack fine-grained distinctions. Performance varies by domain, partly reflecting differences in online visibility. At the language level, accuracy correlates with digital language status: languages with greater digital presence and resources are evaluated more accurately, while low-resource languages perform worse. Analysis of predictive factors confirms that resource-related indicators (Wikipedia size, corpus availability) are more informative than geographic, genealogical, or sociolinguistic factors. Overall, LLM metalinguistic knowledge appears fragmented and shaped mainly by data availability, rather than broadly generalizable grammatical competence. We release the benchmark as an open-source dataset to support evaluation across languages and encourage greater global linguistic diversity in future LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。