arXiv:2503.03962cs.CL2025-03ACL被引 10

探究双语模型如何形成共享语法表征,发现语言差异影响跨语言迁移效果。

On the Acquisition of Shared Grammatical Representations in Bilingual Language Models

  • 通过控制数据量与训练顺序,研究双语模型的语法表征演化过程。
  • 发现语言对间存在不对称的结构启动效应,且相似性越低效果越弱。
  • 为理解多语言模型的共享表征机制提供实验依据,适合关注跨语言学习的研究者。

跨语言迁移对现代语言模型的多语言能力至关重要,但其机制尚不明确。本文探讨单语模型在开始训练第二语言时的变化。我们训练了小型双语模型,精确控制每种语言的数据量和语言暴露顺序。为检验共享多语言表征的存在,采用人类语法表征研究中的结构启动方法。首先复现了先前的跨语言结构启动结果,发现控制数据量和语言暴露后,不同语言对及方向间存在不对称效应。我们认为这种不对称性可能影响对人类结构启动效应的假设。同时发现,语言类型差异越大,结构启动效应越不显著,揭示了跨语言迁移学习与共享表征在类型学多样性语言中的潜在局限。

原文摘要 · Abstract (English)

Crosslingual transfer is crucial to contemporary language models' multilingual capabilities, but how it occurs is not well understood. We ask what happens to a monolingual language model when it begins to be trained on a second language. Specifically, we train small bilingual models for which we control the amount of data for each language and the order of language exposure. To find evidence of shared multilingual representations, we turn to structural priming, a method used to study grammatical representations in humans. We first replicate previous crosslingual structural priming results and find that after controlling for training data quantity and language exposure, there are asymmetrical effects across language pairs and directions. We argue that this asymmetry may shape hypotheses about human structural priming effects. We also find that structural priming effects are less robust for less similar language pairs, highlighting potential limitations of crosslingual transfer learning and shared representations for typologically diverse languages.

双语模型语法表征跨语言迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。