研究语言模型如何在有限数据下学会填补空缺句法关系
Filling in the Mechanisms: How do LMs Learn Filler-Gap Dependencies under Developmental Constraints?
- 用分布式对齐搜索分析模型在不同数据量下的句法表征
- 发现模型在少量数据下仍能发展出共享但敏感的句法机制
- 揭示模型需远超人类的数据量才能达成类似理解,适合语言习得研究者
对于人类而言,填补空缺依赖关系需要跨不同句法结构的共享表征。尽管因果分析表明大语言模型(LLM)也可能存在此类表征(Boguraev等,2025),但在发育可行数据量训练下是否成立尚不清楚。本研究采用分布式对齐搜索(DAS, Geiger等,2024)方法,对在婴儿语言模型挑战(BabyLM, Warstadt等,2023)中以不同数据量训练的语言模型进行分析,评估其在疑问句与话题化结构之间填补空缺依赖的表征是否可迁移。这两类结构在输入频率上差异极大。结果表明,在有限数据下,模型可能发展出共享但对具体词语敏感的机制。更重要的是,模型仍需远超人类的数据量才能掌握可比的泛化能力,凸显了在语言习得模型中引入语言特异性偏置的重要性。
原文摘要 · Abstract (English)
For humans, filler-gap dependencies require a shared representation across different syntactic constructions. Although causal analyses suggest this may also be true for LLMs (Boguraev et al., 2025), it is still unclear if such a representation also exists for language models trained on developmentally feasible quantities of data. We applied Distributed Alignment Search (DAS, Geiger et al. (2024)) to LMs trained on varying amounts of data from the BabyLM challenge (Warstadt et al., 2023), to evaluate whether representations of filler-gap dependencies transfer between wh-questions and topicalization, which greatly vary in terms of their input frequency. Our results suggest shared, yet item-sensitive mechanisms may develop with limited training data. More importantly, LMs still require far more data than humans to learn comparable generalizations, highlighting the need for language-specific biases in models of language acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。