用21种人造语言测试大模型语言学习能力,发现模型更易掌握自然语言衍生的人造语。
ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages

- 构建首个大规模人造语言基准,含21种语言的2100万平行语料
- 模型在源自自然语言的人造语上翻译表现更好,反映设计特征影响学习效果
- 训练可让模型学会8种足够语料的人造语言,适合研究低资源语言学习
人造语言(conlangs)是人类有意创造的语言,具有丰富的语言创造力传统。尽管其在研究大语言模型(LLMs)语言学习方面潜力巨大,但现有研究仍严重忽视了它们的应用。我们提出了ConlangBench,这是首个针对21种现存人造语言的大规模基准,用于评估和训练LLMs。我们收集了超过2100万条人造语言-英语平行句对(其中43万条来自20种非世界语的人造语言),以及32.1万条词汇条目。在双向翻译实验中,我们发现模型在后验型人造语言(a posteriori conlangs)上表现更优,这类语言的词汇源自自然语言,体现了其设计特点。通过在ConlangBench上训练,模型能够学会8种拥有充足平行语料的人造语言,且学习曲线随人造语言的设计方式而异。结果表明,人造语言为探究大模型如何习得低资源语言提供了独特实验平台。
原文摘要 · Abstract (English)
Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for studying language learning in large language models (LLMs), existing conlangs remain largely underexplored in LLM research. We present ConlangBench, the first large-scale benchmark for evaluating and training LLMs on 21 existing conlangs. We collect over 21M conlang-English parallel sentence pairs (including 430K pairs across the 20 non-Esperanto conlangs) and 321K vocabulary entries. In bidirectional translation experiments, we find that models perform better on a posteriori conlangs, whose vocabularies are derived from natural languages, reflecting the design characteristics of conlangs. Training on ConlangBench also shows that models can learn all eight conlangs for which sufficient parallel corpora are available, while their learning curves vary depending on how the conlangs were created. Our findings suggest that conlangs provide a unique testbed for investigating how LLMs acquire low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。