人工神经网络训练中自发出现语音、词汇和语法表征,揭示语言学习的计算机制。
Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks
- 通过分析网络激活态的几何结构,发现语音、词汇、语法表征逐级涌现。
- 语音与词汇表征在训练初期形成,语法表征需更多数据,约需100~1000倍数据量。
- 结果为理解语言习得的神经机制提供可验证的计算框架,适合认知科学与AI交叉研究者。
在语言习得过程中,儿童依次掌握音素分类、词汇识别及语法组合生成新意义。尽管这一行为发展路径已被充分描述,但缺乏统一的计算框架来解释其背后的神经表征。本文研究了人工神经网络在训练过程中何时以及如何涌现出音素、词汇和句法表征。结果显示,基于语音和文本的模型均遵循学习阶段序列:随着训练推进,其神经激活逐渐构建出子空间,其中激活态的几何结构分别对应音素、词汇和句法结构。该发展轨迹在定性上与儿童语言习得相似,但在定量上存在差异:这些算法需要比人类多出两到四个数量级的数据才能使表征浮现。这些发现明确了语言习得关键阶段自发出现的条件,为理解语言习得背后计算机制指明了可行路径。
原文摘要 · Abstract (English)
During language acquisition, children successively learn to categorize phonemes, identify words, and combine them with syntax to form new meaning. While the development of this behavior is well characterized, we still lack a unifying computational framework to explain its underlying neural representations. Here, we investigate whether and when phonemic, lexical, and syntactic representations emerge in the activations of artificial neural networks during their training. Our results show that both speech- and text-based models follow a sequence of learning stages: during training, their neural activations successively build subspaces, where the geometry of the neural activations represents phonemic, lexical, and syntactic structure. While this developmental trajectory qualitatively relates to children's, it is quantitatively different: These algorithms indeed require two to four orders of magnitude more data for these neural representations to emerge. Together, these results show conditions under which major stages of language acquisition spontaneously emerge, and hence delineate a promising path to understand the computations underpinning language acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。