arXiv:2503.20850cs.CL2025-03被引 14

语言模型的语法偏好源于直接证据与间接规律的共同作用。

Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models

  • 通过控制输入数据,分离直接与间接语言规律的影响。
  • 即使无直接证据,模型仍倾向短语在前、有生命者在前。
  • 适合研究语言模型语法习得机制的研究者阅读。

语言模型(LMs)在多项句法现象上表现出类人偏好,但这些偏好是源于对具体现象的直接接触,还是更普遍的语言规律尚不明确。本文以英语双宾结构(如:给某人某物 / 给某物给某人)为例,采用受控训练范式,迭代训练小型语言模型于系统性改写的输入数据。重点关注影响选择的因素:长度与生命性。这两者既在双宾结构中直接体现,也反映更广泛的语言倾向——短成分优先于长成分,有生命者优先于无生命者。通过操纵并消除输入中的长度与生命性偏差,我们发现直接证据确实影响模型选择,但“先易后难”偏好仍在无直接证据时持续存在。进一步通过重构数据集,全局重排句子顺序而保留依存结构,我们发现双宾偏好仍可由间接语言规律引发。结论表明,语言模型的句法偏好源自直接与间接证据的混合机制。

原文摘要 · Abstract (English)

Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general properties of language is unclear. We explore this with the English dative alternation (DO: "gave Y the X" vs. PO: "gave the X to Y"), using a controlled rearing paradigm wherein we iteratively train small LMs on systematically manipulated input. We focus on two properties that affect the choice of alternant: length and animacy. Both properties are directly present in datives but also reflect more global tendencies for shorter elements to precede longer ones and animates to precede inanimates. First, by manipulating and ablating datives for these biases in the input, we show that direct evidence of length and animacy matters, but easy-first preferences persist even without such evidence. Then, using LMs trained on systematically perturbed datasets to manipulate global length effects (re-linearizing sentences globally while preserving dependency structure), we find that dative preferences can emerge from indirect evidence. We conclude that LMs' emergent syntactic preferences come from a mix of direct and indirect sources.

语言模型句法偏好双宾结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。