用上下文统计提升儿童语言习得中的单词分割能力
Using Context to Improve Word Segmentation
- 对比一元与二元统计模型,利用上下文信息改进分割
- 二元模型在预测单词边界上显著优于一元模型
- 模拟儿童用已学词汇辅助新词分割,适合语言认知研究
理解儿童语言习得的关键步骤是研究婴儿如何进行单词分割。已有研究表明,婴儿可能通过语音中的统计规律学习单词分割。Goldwater 等人的研究证明,在模型中引入上下文能提升其学习效果。我们实现了他们的一个一元模型和一个二元模型,以检验上下文对统计单词分割的改善作用。结果支持假设:二元模型在预测单词分割方面优于一元模型。此外,我们还探索了年幼儿童如何利用已学词汇来分割新话语的基本建模方式。
原文摘要 · Abstract (English)
An important step in understanding how children acquire languages is studying how infants learn word segmentation. It has been established in previous research that infants may use statistical regularities in speech to learn word segmentation. The research of Goldwater et al., demonstrated that incorporating context in models improves their ability to learn word segmentation. We implemented two of their models, a unigram and bigram model, to examine how context can improve statistical word segmentation. The results are consistent with our hypothesis that the bigram model outperforms the unigram model at predicting word segmentation. Extending the work of Goldwater et al., we also explored basic ways to model how young children might use previously learned words to segment new utterances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。