arXiv:2603.08359cs.CLcs.AI2026-03被引 3

无需语言先验,模型通过听觉和视听输入学习语言,揭示早期语言发展的共性机制。

Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors

  • 使用自监督与视觉引导的感知学习模型模拟婴儿语言习得
  • 模型在无语言先验下成功习得语音特征,符合婴儿发展规律
  • 适用于研究语言发展、认知科学及通用学习理论的跨学科研究

理解语音对正常发育的婴幼儿而言似乎毫不费力,但从信息处理角度看,仅凭声音信号习得语言是一项巨大挑战。本文综述了近年来利用计算模型理解从语音与视听输入中早期语言习得的进展。重点聚焦于自监督与视觉引导的感知学习模型。研究表明,这些模型在无需强语言先验条件下,日益强大地学习语音的多个层面;同时,许多早期语言发展的特征可通过一组共享的学习原则解释,这些原则与多种语言习得与人类认知理论相兼容。此外,现代学习模拟正逐步提升真实性,体现在输入数据的逼真性以及模型行为与婴儿语言发展实证发现的关联上。

原文摘要 · Abstract (English)

Learning to understand speech appears almost effortless for typically developing infants, yet from an information-processing perspective, acquiring a language from acoustic speech is an enormous challenge. This chapter reviews recent developments in using computational models to understand early language acquisition from speech and audiovisual input. The focus is on self-supervised and visually grounded models of perceptual learning. We show how these models are becoming increasingly powerful in learning various aspects of speech without strong linguistic priors, and how many features of early language development can be explained through a shared set of learning principles-principles broadly compatible with multiple theories of language acquisition and human cognition. We also discuss how modern learning simulations are gradually becoming more realistic, both in terms of input data and in linking model behavior to empirical findings on infant language development.

语言习得自监督学习认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。