用扰动提升大模型对训练外文本的预测能力
Perturbation is All You Need for Extrapolating Language Models

- 通过引入前缀扰动生成语义邻近变体,再基于扰动后输入预测下一个词
- 在合成与真实语料上均显著提升训练外序列的预测性能
- 适合需要泛化到罕见或未见语言模式的研究者
本文构建了大语言模型外推能力的统计理论,通过重新诠释为预-后加性噪声模型。不同于传统的精确前缀自回归预测,本文提出一种基于扰动的方法:先将前缀转换为语义邻近变体,再以此扰动版本进行下一步词预测。该方法形成具有预-后加性噪声结构的分层模型。在此框架下,建立了外推可实现性的严格理论,证明了所提方法具备自适应性、收缩性、鲁棒性、外推能力与双重鲁棒性五项性质。在合成与真实语言数据上的有限样本实验表明,该方法持续提升训练外序列的预测表现,同时保持与现有方法相当的训练内性能,验证了扰动是实现语言建模外推的可行路径。
原文摘要 · Abstract (English)
This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models. In contrast to the standard autoregressive next-token prediction based on an exact prefix, we introduce a perturbation-based procedure that first transforms the prefix into a semantic neighbour and then conditions on this perturbed variant for next-token prediction. This yields a hierarchical model with a pre-post-additive noise structure. Within this framework, we develop a rigorous theory of extrapolability, namely, the capacity of a model class to make reliable predictions for token sequences that lie outside the empirical support of the training corpus, by establishing five properties of the proposed procedure: adaptivity, contractivity, robustness, extrapolability, and double robustness. We evaluate the finite sample performance of the proposed procedure using both synthetic and real world language data. Results show that the proposed method consistently improves out-of-support prediction while maintaining competitive in-support performance, demonstrating that perturbation offers a practical route to language modelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。