通过替换组件提升语言模型的组合泛化能力,解决复杂语义组合难题。
Learning to Substitute Components for Compositional Generalization
- 提出组件替换策略,实现跨层级的复合结构增强
- 在多个基准上提升组合泛化性能,最高达66.5%改善
- 适用于大模型少样本学习场景,适合需要强泛化能力的研究者
尽管神经语言模型日益普及,但近期实证表明其在组合泛化方面存在缺陷。当前主流解决方案是组合数据增强,旨在引入额外的组合归纳偏置。然而,现有手工设计的增强策略在系统性泛化需要多粒度组合偏置(即不仅限于词汇或结构偏置)或训练句难度分布不均时,改进有限。为此,我们首先提出一种新型组合增强策略——组件替换(CompSub),可在整个训练集中实现大量子结构的多粒度组合。进一步提出学习组件替换(LCS)框架,通过最大化语言模型损失,端到端学习组件替换概率,优先关注具有隐蔽概念和新语境的困难组合。我们将CompSub和LCS的核心思想扩展至预训练大模型的上下文学习场景,提出LCS-ICL算法,以增强当前最优大模型的少样本组合泛化能力。理论上,我们解释了为何该算法能提升语言模型的组合泛化表现。实验结果在四个标准组合泛化基准(SCAN、COGS、GeoQuery、COGS-QL)上显示,CompSub、LCS、LCS-ICL分别取得最高66.5%、10.3%、1.4%和8.8%的性能提升。
原文摘要 · Abstract (English)
Despite the rising prevalence of neural language models, recent empirical evidence suggests their deficiency in compositional generalization. One of the current de-facto solutions to this problem is compositional data augmentation, which aims to introduce additional compositional inductive bias. However, existing handcrafted augmentation strategies offer limited improvement when systematic generalization of neural language models requires multi-grained compositional bias (i.e., not limited to either lexical or structural biases alone) or when training sentences have an imbalanced difficulty distribution. To address these challenges, we first propose a novel compositional augmentation strategy called Component Substitution (CompSub), which enables multi-grained composition of substantial substructures across the entire training set. Furthermore, we introduce the Learning Component Substitution (LCS) framework. This framework empowers the learning of component substitution probabilities in CompSub in an end-to-end manner by maximizing the loss of neural language models, thereby prioritizing challenging compositions with elusive concepts and novel contexts. We extend the key ideas of CompSub and LCS to the recently emerging in-context learning scenarios of pre-trained large language models (LLMs), proposing the LCS-ICL algorithm to enhance the few-shot compositional generalization of state-of-the-art (SOTA) LLMs. Theoretically, we provide insights into why applying our algorithms to language models can improve compositional generalization performance. Empirically, our results on four standard compositional generalization benchmarks(SCAN, COGS, GeoQuery, and COGS-QL) demonstrate the superiority of CompSub, LCS, and LCS-ICL, with improvements of up to 66.5%, 10.3%, 1.4%, and 8.8%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。