arXiv:2608.25976cs.CLcs.LG2026-08

用神经模型模拟被收养者,发现母语痕迹存于底层,重学更快。

Lost but not erased: Finding traces of a forgotten language in neural speech models

  • 用语音识别模型模拟断续语言学习,无生理成熟干扰。
  • 第一语言痕迹保留在低层网络,重学速度提升14%。
  • 适合研究语言习得机制与神经网络表征的读者。

国际收养者虽无法再使用或理解出生语言,但仍保留其语音痕迹,传统观点认为这源于生物性关键期。我们通过自动语音识别模型模拟收养经历,排除发育因素干扰:模型先训练一种语言,再突然切换至第二种。结果发现,第一语言痕迹在整个第二语言训练中持续存在,主要集中在最低、前音素层级。这些痕迹具有功能性:有早期语言经验的模型重学第一语言的速度比从未接触过的模型快14%;即使对比从相关语言早期收养的模型,该优势依然存在;但若替换最底层参数为未收养模型,则优势消失。我们提出,关键期效应并非源于可塑性丧失,而是基础表征固化所致,经验在语言关键期中起核心作用。

原文摘要 · Abstract (English)

International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the ordinary dynamics of learning, using automatic speech recognition models that simulate the international adoptee experience without maturational confounds. Models were trained on one language and then abruptly switched to a second. We found that traces of the first language persisted throughout second-language training, but mainly in the lowest, pre-phonemic layers. These traces were functional, as models with early exposure re-learned their lost first language 14% faster than naive models; this advantage held even against models adopted early from a related language and disappeared when the earliest layers were substituted from a non-adopted model. We argue that these critical-period effects reflect entrenchment of foundational representations rather than a maturational loss of plasticity, and that experience plays a central role in critical periods in language acquisition.

语言习得神经模型表征固化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。