arXiv:2603.19427cs.CLcs.AI2026-03ACL

词汇结构决定语言模型对语序学习的难易程度

Vocabulary shapes cross-lingual variation of word-order learnability in language models

  • 用合成语序变异训练模型,测试不同语言的可学性
  • 词汇越不规则,模型困惑度越高,学习越困难
  • 适合研究语言差异与大模型语言能力的读者

为什么一些语言如捷克语允许自由语序,而英语等则不行?我们通过在自然语言的合成语序变体上预训练Transformer语言模型来回答这一问题。发现语序不规则性越高,模型困惑度(surprisal)越大,表明学习难度增加。句子倒序对学习的影响较弱。仅区分自由语序(如捷克语、芬兰语)与固定语序(如英语、法语)无法解释跨语言差异。相反,词和子词词汇结构能有效预测模型困惑度。总体而言,词汇结构是跨语言语序学习性的关键驱动因素。

原文摘要 · Abstract (English)

Why do some languages like Czech permit free word order, while others like English do not? We address this question by pretraining transformer language models on a spectrum of synthetic word-order variants of natural languages. We observe that greater word-order irregularity consistently raises model surprisal, indicating reduced learnability. Sentence reversal, however, affects learnability only weakly. A coarse distinction of free- (e.g., Czech and Finnish) and fixed-word-order languages (e.g., English and French) does not explain cross-lingual variation. Instead, the structure of the word and subword vocabulary strongly predicts the model surprisal. Overall, vocabulary structure emerges as a key driver of computational word-order learnability across languages.

语言模型语序学习词汇结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。