研究语言模型中双语学习的结构干扰机制,揭示语言主导与熟练度如何影响跨语言影响。
A Study of Crosslinguistic Influence in Language Models
- 通过控制第二语言引入时机,模拟顺序双语学习过程
- 高母语主导性增强结构迁移,但低第二语言熟练度限制正向迁移
- 深层注意力机制负责跨语言结构传递,且受母语类型距离影响
语言的顺序习得必然导致跨语言影响(CLI),即第一语言(L1)的句法特性影响第二语言(L2)的处理。尽管现代语言模型具备强大的跨语言迁移能力,但语言主导性、相对熟练度与类型学距离如何调控结构干扰仍需深入探究。本文系统研究人工双语学习者中的CLI,通过在15种类型多样的L1上训练同时与顺序双语模型,并改变语言引入步骤(SoE),利用跨语言结构启动技术,将隐含的CLI解耦为正向与负向迁移率。评估显示:更高的L1主导性(高SoE)显著增强结构迁移与句法距离的相关性,但较低的L2熟练度会抑制正向迁移,使模型极易受到远距离L1的持续负向干扰。机制层面,我们证明L1类型相近性决定了L2句法神经元的跨语言重叠;此外,显式引导引发语法解析在层间动态迁移至网络末端。通过定向因果消融验证,仅深层注意力机制驱动跨语言迁移,通过将L1结构先验导向最终预测。结果表明,语言模型中的CLI并非容量限制的偶然产物,而是由语言主导性与熟练度相互作用决定的结构性现象。
原文摘要 · Abstract (English)
The sequential acquisition of languages inevitably leads to Crosslinguistic Influence (CLI), where the syntactic properties of a first language (L1) impact the processing of a second language (L2). While modern language models exhibit robust cross-lingual transfer, the exact mechanisms governing how language dominance, relative proficiency, and typological distance dictate structural interference warrant deeper investigation. In this work, we systematically investigate CLI in artificial learners by training simultaneous and sequential bilingual models across 15 typologically diverse L1s and varying the Step of Exposure (SoE), defined as the specific training step at which the L2 is introduced. Utilizing crosslinguistic structural priming, we decouple latent CLI into distinct positive and negative transfer rates. Our evaluations reveal a critical computational tradeoff: while increased L1 dominance (higher SoE) strongly amplifies the correlation between structural transfer and syntactic distance, diminished L2 proficiency bottlenecks the model's capacity for positive transfer, leaving it highly vulnerable to persistent negative interference from distant L1s. Mechanistically, we demonstrate that L1 typological proximity physically dictates the cross-lingual overlap of L2 syntactic neurons. Furthermore, we uncover that explicit priming induces a dynamic layer-wise migration of grammatical resolution to the terminal layers of the network. Through targeted causal ablations, we confirm that deep-layer attention mechanisms exclusively drive this crosslinguistic transfer by routing the L1 structural prior into the final prediction. Ultimately, our findings demonstrate that CLI in language models is not an arbitrary artifact of capacity constraints, but a structured phenomenon fundamentally governed by the interplay of language dominance and proficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。