揭示语言模型如何逐层构建语法结构,发现低层生成局部语法、高层整合全局语法。
Derivational Probing: Unveiling the Layer-wise Derivation of Syntactic Structures in Neural Language Models
- 通过层级探针分析语法结构在模型中的逐步生成过程。
- Bert中局部语法结构在底层出现,高层形成完整句法关系。
- 句法信息整合时机影响下游任务性能,存在最优整合时间点。
近期研究表明神经语言模型在其内部表示中编码了句法结构,但这些结构如何在各层中逐步构建仍不清楚。本文提出衍生探针(Derivational Probing),研究微语法结构(如主语名词短语)和宏语法结构(如动词与其直接依赖项的关系)如何随词向量在层间传播而逐步形成。在BERT上的实验表明,存在清晰的自底向上构建路径:微语法结构在较低层出现,并逐渐整合为高层的一致性宏语法结构。此外,针对主谓数一致性的针对性评估显示,宏语法结构的构建时机对下游性能至关重要,提示全局句法信息应存在最优整合时间。
原文摘要 · Abstract (English)
Recent work has demonstrated that neural language models encode syntactic structures in their internal representations, yet the derivations by which these structures are constructed across layers remain poorly understood. In this paper, we propose Derivational Probing to investigate how micro-syntactic structures (e.g., subject noun phrases) and macro-syntactic structures (e.g., the relationship between the root verbs and their direct dependents) are constructed as word embeddings propagate upward across layers. Our experiments on BERT reveal a clear bottom-up derivation: micro-syntactic structures emerge in lower layers and are gradually integrated into a coherent macro-syntactic structure in higher layers. Furthermore, a targeted evaluation on subject-verb number agreement shows that the timing of constructing macro-syntactic structures is critical for downstream performance, suggesting an optimal timing for integrating global syntactic information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。