arXiv:2511.00198cs.CLcs.AI2025-11被引 2

提出新训练方法,突破传统逐词预测局限

Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap

  • 用信息量高的目标词替代逐词预测,优化训练策略
  • 在算术、文本多标签分类和生成任务中提升模型表现
  • 适合关注训练效率与理论机制的LLM研究者

大语言模型(LLM)的训练性能优化仍是关键挑战,尤其在保持计算成本的同时提升模型能力。本文质疑传统基于下一个词预测(NTP)的训练方式,认为通过在训练中预测信息量更高的目标词,可实现更高效的训练。研究在三类任务上验证该方法:算术推理、文本多标签分类和自然语言生成。结果表明,该方法能有效提升模型性能,并深化对目标词选择策略的理论理解,为优化大模型训练提供了系统性思路。

原文摘要 · Abstract (English)

Optimizing training performance in large language models (LLMs) remains an essential challenge, particularly in improving model performance while maintaining computational costs. This work challenges the conventional approach of training LLMs using next-token prediction (NTP), arguing that by predicting information-rich tokens during training, there is a more effective way to train LLMs. We investigate the impact of the proposed solution in three kinds of tasks for LLMs: arithmetic, multi-label classification of text, and natural-language generation. This work offers a principled approach to optimizing LLM training, advancing both model performance and theoretical understanding of the target-token selection strategies.

大模型训练信息最大化目标词选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。