用真实儿童语言数据训练模型,学会远距离语法依赖的解析。
Modelling Child Learning and Parsing of Long-range Syntactic Dependencies
- 基于真实儿童语料,同时学习词义与语法结构
- 能准确推断语句含义和语法树,支持单语句意义推理
- 突破上下文无关语法限制,对语言习得理论有启发
本文构建了一个概率性儿童语言习得模型,用于学习多种语言现象,尤其是对象疑问句等结构中的远距离句法依赖。模型在真实儿童导向言语语料库上训练,每个话语都配有逻辑形式作为语义表示。训练后,模型可同时推断给定话语-语义对的正确语法树和词义,并在仅提供话语时反推其语义。成功建模远距离依赖在理论上具有重要意义,因其利用了通常超出上下文无关语法范围的建模能力。
原文摘要 · Abstract (English)
This work develops a probabilistic child language acquisition model to learn a range of linguistic phenonmena, most notably long-range syntactic dependencies of the sort found in object wh-questions, among other constructions. The model is trained on a corpus of real child-directed speech, where each utterance is paired with a logical form as a meaning representation. It then learns both word meanings and language-specific syntax simultaneously. After training, the model can deduce the correct parse tree and word meanings for a given utterance-meaning pair, and can infer the meaning if given only the utterance. The successful modelling of long-range dependencies is theoretically important because it exploits aspects of the model that are, in general, trans-context-free.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。