通过自动识别儿童语言输入中的空缺依赖结构,揭示其语言习得的真实语料基础。
What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies
- 结合成分与依存句法分析,自动识别三种核心空缺结构及提取位置。
- 在57个英文儿童语料库中发现不同构造的出现频率与主宾位不对称性。
- 为语言习得与语言模型训练提供可量化的细粒度输入数据,适合研究者使用。
儿童习得空缺依赖结构是否依赖先天语法知识,还是仅靠儿童导向话语中的分布证据即可解释,这一问题长期存在争议。然而,相关输入难以大规模、精细量化,阻碍了研究进展。本文提出一种系统,可在口语语料库中自动识别三种核心空缺结构:主句疑问句、嵌套疑问句和关系从句,并进一步判断提取位置(主语、宾语或状语)。该方法融合成分句法与依存句法分析,发挥二者互补优势。在人工标注数据上验证,系统在多数类别表现良好。应用于57个英文CHILDES语料库后,我们刻画了儿童在发展过程中对各类空缺结构的输入频率及提取位置偏差,提供了细粒度标签。这些结果可支持未来习得与计算研究,文中以过滤语料训练语言模型为例进行了演示。
原文摘要 · Abstract (English)
Children's acquisition of filler-gap dependencies has been argued by some to depend on innate grammatical knowledge, while others suggest that the distributional evidence available in child-directed speech suffices. Unfortunately, the relevant input is difficult to quantify at scale with fine granularity, making this question difficult to resolve. We present a system that identifies three core filler-gap constructions in spoken English corpora -- matrix wh-questions, embedded wh-questions, and relative clauses -- and further identifies the extraction site (i.e., subject vs. object vs. adjunct). Our approach combines constituency and dependency parsing, leveraging their complementary strengths for construction classification and extraction site identification. We validate the system on human-annotated data and find that it scores well across most categories. Applying the system to 57 English CHILDES corpora, we are able to characterize children's filler-gap input and their filler-gap production trajectories over the course of development, including construction-specific frequencies and extraction-site asymmetries. The resulting fine-grained labels enable future work in both acquisition and computational studies, which we demonstrate with a case study using filtered corpus training with language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。