arXiv:2601.03448cs.CL2026-01ACL被引 3

用语言学习任务预训练,让大模型更懂语法和语言结构。

Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks

  • 将原始文本转为输入输出对,模拟人类学语言过程
  • 在语法任务上表现提升,且学习速度更快
  • 适合想提升模型语言准确性的研究者

语言模型(LMs)通常在原始文本数据上进行预训练,以逐词生成文本序列。这种方法虽有助于学习世界知识与推理能力,但未显式优化语言能力。为弥补这一差距,我们提出L2T框架,在标准的下一个词预测之外引入语言学习任务。受人类语言习得启发,L2T将原始文本转化为结构化的输入-输出对,提供明确的语言刺激。在原始文本与L2T数据混合的预训练下,语言模型不仅在语言能力基准测试中表现更优,且语言能力获取速度加快,同时保持在通用推理任务上的竞争力。

原文摘要 · Abstract (English)

Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token. While this approach facilitates the learning of world knowledge and reasoning, it does not explicitly optimize for linguistic competence. To bridge this gap, we propose L2T, a pre-training framework integrating Language Learning Tasks alongside standard next-token prediction. Inspired by human language acquisition, L2T transforms raw text into structured input-output pairs to provide explicit linguistic stimulation. Pre-training LMs on a mixture of raw text and L2T data not only improves overall performance on linguistic competence benchmarks but accelerates its acquisition, while maintaining competitive performance on general reasoning tasks.

语言模型预训练语法能力学习任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。