用形式推导预训练,让大模型更快学会语言技能并更易压缩。
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

- 在形式推导数据上预训练,引入更强的结构化语言先验。
- 仅用360亿更少的词元就达到80%语言任务准确率,优于传统初始化。
- 模型内部表征更紧凑,剪枝至约33%稀疏度仍保持完整性能。
在自然语言学习前,对语言模型进行符号数据上的预预训练可加速并提升其表现。然而,现有预预训练任务(如Dyck和过程算法)依赖狭窄的原始操作,难以捕捉自然语言的表达能力;且以往研究受限于较小的词元预算,难以揭示技能涌现与表征动态。为此,我们提出逻辑预预训练(Logic-PPT),通过形式推导作为系统性初始化策略,注入更丰富的结构与语言偏置。形式推导需抽象机制,涵盖变量绑定、量词与关系依赖连接,以及长程上下文中的谓词-论元结构组合。将评估规模扩展至1000亿词元,逻辑预预训练显著加速模型技能获取:在语言任务上以360亿更少词元达到80%准确率,超越其他预预训练基线。机制层面,形式推导引发持久的结构重排,表现为更低秩、谱集中度更高的表征空间。关键的是,该内部几何结构使模型可通过剪枝实现更好压缩,在约33%稀疏度下仍匹配稠密基线性能。
原文摘要 · Abstract (English)
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining tasks, such as Dyck and procedural algorithms, rely on narrow primitives that fail to capture the expressive capacity of natural language. Moreover, prior studies remain restricted to relatively small token budgets, offering limited insight into skill emergence and representational dynamics. To address these limitations, we propose logic pre-pretraining (Logic-PPT) as a principled initialization strategy, leveraging formal derivations to impart richer structural and linguistic biases. Formal derivations require abstract mechanisms that are central to natural language, simultaneously binding variables, connecting quantifiers and relational dependencies, and composing predicate-argument structures over long contexts. Scaling our evaluation to a 100B-token regime, logic pre-pretraining substantially accelerates skill acquisition in LMs, achieving 80\% accuracy on linguistic tasks with 36B fewer tokens than standard initialization, and outperforming alternative pre-pretraining baselines. Mechanistically, formal derivations induce persistent structural reorganization, distinctively characterized by a lower-rank, spectrally concentrated representation space. Crucially, we show that this internal geometry enables improved model compressibility via pruning, matching the dense baseline performance even at $\approx$33\% sparsity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。