研究词序对语言模型长句泛化能力的影响,发现符合语言类型学的词序更易泛化。
Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
- 用广义范畴语法构建新型人工语言,覆盖自然语言中未被充分考虑的复杂结构
- 在更长未见句子上测试模型,发现典型词序的泛化性能显著更好
- 为理解语言模型的归纳偏置提供新实验范式,适合关注模型泛化能力的研究者
语言模型是否倾向于偏好语言类型学上常见的语法特征而非罕见、不合理的特征,已有研究多通过人工语言(ALs)进行探讨(White and Cotterell, 2021;Kuribayashi et al., 2024)。本文从两个角度拓展了这些工作:首先,将原有的上下文无关人工语言形式化扩展为广义范畴语法(GCG)(Wood, 2014),使人工语言能够涵盖此前被忽视但实际存在的语言结构,如无界依存和弱上下文敏感结构;其次,评估重点转向语言模型对未见过的更长测试句的泛化能力。因此,本研究构建的人工语言更贴近自然语言特征,实验范式也得出更清晰的结论——语言类型学上合理的词序更易于让语言模型实现有效的泛化。
原文摘要 · Abstract (English)
Whether language models (LMs) have inductive biases that favor typologically frequent grammatical properties over rare, implausible ones has been investigated, typically using artificial languages (ALs) (White and Cotterell, 2021; Kuribayashi et al., 2024). In this paper, we extend these works from two perspectives. First, we extend their context-free AL formalization by adopting Generalized Categorial Grammar (GCG) (Wood, 2014), which allows ALs to cover attested but previously overlooked constructions, such as unbounded dependency and mildly context-sensitive structures. Second, our evaluation focuses more on the generalization ability of LMs to process unseen longer test sentences. Thus, our ALs better capture features of natural languages and our experimental paradigm leads to clearer conclusions -- typologically plausible word orders tend to be easier for LMs to productively generalize.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。