arXiv:2602.17653cs.CL2026-02被引 1

语言模型在论元标记中表现出与人类相似的标记方向偏好,但不复制对象优先现象。

Differences in Typological Alignment in Language Models' Treatment of Differential Argument Marking

  • 用合成数据训练GPT-2模型,模拟18种不同论元标记系统。
  • 模型正确偏好语义非典型论元被显式标记,但未复现人类的对象优先倾向。
  • 揭示不同类型语言规律可能源于不同认知机制,适合研究语言习得与模型偏差。

近期研究表明,基于合成语料训练的语言模型会表现出类似人类语言跨语言规律的类型学偏好,尤其体现在词序等句法现象上。本文将此范式扩展至差异性论元标记(DAM),一种依赖语义突出性的形态标记系统。我们采用受控的合成学习方法,训练GPT-2模型在18个实现不同DAM系统的语料上,并通过最小对立对评估其泛化能力。结果揭示了DAM两种类型学维度的分离:模型可靠地表现出人类类似的自然标记方向偏好,即更倾向于显式标记语义非典型的论元;然而,模型未能再现人类语言中强烈的对象偏好——即显式标记更常作用于宾语而非主语。这些发现表明,不同类型的语言规律可能源自不同的内在机制。

原文摘要 · Abstract (English)

Recent work has shown that language models (LMs) trained on synthetic corpora can exhibit typological preferences that resemble cross-linguistic regularities in human languages, particularly for syntactic phenomena such as word order. In this paper, we extend this paradigm to differential argument marking (DAM), a semantic licensing system in which morphological marking depends on semantic prominence. Using a controlled synthetic learning method, we train GPT-2 models on 18 corpora implementing distinct DAM systems and evaluate their generalization using minimal pairs. Our results reveal a dissociation between two typological dimensions of DAM. Models reliably exhibit human-like preferences for natural markedness direction, favoring systems in which overt marking targets semantically atypical arguments. In contrast, models do not reproduce the strong object preference in human languages, in which overt marking in DAM more often targets objects rather than subjects. These findings suggest that different typological tendencies may arise from distinct underlying sources.

语言模型论元标记类型学生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。