arXiv:2410.18225cs.CL2024-10中稿 · CoNLL 2024被引 7

神经语言模型难以捕捉填充空位依赖的深层结构共性。

Generalizations across filler-gap dependencies in neural language models

  • 通过控制输入,测试模型是否能识别填充空位的共性结构
  • 模型仅依赖表面特征区分语法正确与错误句式
  • 研究提示需加入语言习得的特殊先验知识

人类通过有限输入发展语法,基于结构上的泛化。我们探究填充空位依赖(尽管表面形式多样)如何从输入中产生。通过显式控制神经语言模型(NLM)的输入,考察模型是否对这类依赖建立共享表征。结果表明,尽管模型能成功区分语法正确与错误的填充空位句,但其判断依赖于输入的表面属性,而非共享的深层结构泛化。该研究强调了在建模语言习得时引入特定语言学归纳偏置的必要性。

原文摘要 · Abstract (English)

Humans develop their grammars by making structural generalizations from finite input. We ask how filler-gap dependencies, which share a structural generalization despite diverse surface forms, might arise from the input. We explicitly control the input to a neural language model (NLM) to uncover whether the model posits a shared representation for filler-gap dependencies. We show that while NLMs do have success differentiating grammatical from ungrammatical filler-gap dependencies, they rely on superficial properties of the input, rather than on a shared generalization. Our work highlights the need for specific linguistic inductive biases to model language acquisition.

语言模型语法泛化结构共性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。