探究结构诱导语言模型的语法泛化能力,发现新模型在长距离依赖上表现更优。
Understanding Syntactic Generalization in Structure-inducing Language Models
- 对比三种结构诱导模型的语法表示生成机制
- GPST模型在合成括号表达式中对长距离依赖建模最优
- 大规模合成数据小模型可作为基础性质测试平台
结构诱导语言模型(SiLM)通过自监督语言建模训练,在处理输入时会生成层次化句子表示。尽管其具备强语法泛化能力且在多种NLP任务中表现优异,但其基本属性仍不明确。本文从头训练了三种不同架构:Structformer(Shen et al., 2021)、UDGN(Shen et al., 2022)和GPST(Hu et al., 2024b),分别在自然语言(英语、德语、中文)和合成括号表达式上进行训练。评估涵盖三方面:诱导的句法表示特性、语法判断任务表现及训练动态。结果表明,三种模型无一在所有指标上全面领先,但在句法表示方面存在显著差异。其中,生成式预训练结构变换器(GPST)在各类设置中表现最稳定,并在括号表达式的长距离依赖任务中优于其他模型。研究还发现,用大量合成数据训练的小模型可有效用于评估模型基础性质。
原文摘要 · Abstract (English)
Structure-inducing Language Models (SiLM) are trained on a self-supervised language modeling task, and induce a hierarchical sentence representation as a byproduct when processing an input. SiLMs couple strong syntactic generalization behavior with competitive performance on various NLP tasks, but many of their basic properties are yet underexplored. In this work, we train three different SiLM architectures from scratch: Structformer (Shen et al., 2021), UDGN (Shen et al., 2022), and GPST (Hu et al., 2024b). We train these architectures on both natural language (English, German, and Chinese) corpora and synthetic bracketing expressions. The models are then evaluated with respect to (i) properties of the induced syntactic representations (ii) performance on grammaticality judgment tasks, and (iii) training dynamics. We find that none of the three architectures dominates across all evaluation metrics. However, there are significant differences, in particular with respect to the induced syntactic representations. The Generative Pretrained Structured Transformer (GPST; Hu et al. 2024) performs most consistently across evaluation settings, and outperforms the other models on long-distance dependencies in bracketing expressions. Furthermore, our study shows that small models trained on large amounts of synthetic data provide a useful testbed for evaluating basic model properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。