arXiv:2608.05028cs.CL2026-08

语言模型在缺乏直接证据时仍偏好人类的词序习惯。

Language Models Generalize to Human-like Word Order Preferences

论文配图:Language Models Generalize to Human-like Word Order Preferences
图 1 · 摘自论文原文
  • 用移除多修饰语短语的语料训练模型,测试其词序偏好。
  • 三种规模模型均倾向选择符合语义范围一致性的词序。
  • 发现该偏好非源于词语关联强度,揭示潜在学习机制。

语言习得中的核心问题是:语言偏见能否从对不充分输入进行通用学习的过程中产生?人工语言学习研究显示,人类学习者能可靠地超越已有证据进行泛化,包括偏好语义范围同构的名词短语修饰语顺序。本文研究语言模型是否在类似条件下也表现出这种偏好。我们构建了一个受控学习环境:在训练语料中移除所有含多个修饰语的名词短语,从而消除关于修饰语顺序的直接证据,随后在包含多修饰语的句子上评估模型表现。在三个不同规模的语言模型中,我们发现它们尽管从未在训练中见过此类结构,却仍持续偏好语义范围同构的修饰语顺序,且偏好强度随修饰语类型而异。为探究该偏好的来源,我们使用点互信息(PMI)分析名词与修饰语间的关联强度,结果表明:虽PMI反映已知的修饰语顺序模式,但无法解释模型的排序偏好。这些发现表明,语言模型可从贫乏输入中恢复类人的语言泛化能力,并为研究此类偏见的形成机制提供了可控框架。

原文摘要 · Abstract (English)

A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input. Artificial Language Learning (ALL) studies have shown that human learners reliably generalize beyond the evidence provided, including by preferring scope-homomorphic noun phrase modifier orders. In this work, we investigate whether language models exhibit the same bias under similar conditions. We create a controlled learning environment in which models are trained on a corpus where all noun phrases containing multiple modifiers have been removed, eliminating direct evidence about modifier ordering, and are then evaluated on multiple modifier sentences. Across three model sizes, we find that they consistently prefer scope-homomorphic orders despite never observing them during training. These preferences vary in strength by modifier type. To investigate the source of these preferences, we examine noun-modifier association strength using pointwise mutual information (PMI). While PMI reflects known modifier-ordering patterns, it does not explain the models' ordering preferences. These findings demonstrate that LMs can recover human-like linguistic generalizations from impoverished input and provide a controlled framework for investigating the mechanisms underlying such biases.

语言模型词序偏好泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。