arXiv:2509.24435cs.CLcs.AI2025-09综述被引 1

探索取代逐词生成的新范式,突破大模型的生成瓶颈

Alternatives To Next Token Prediction In Text Generation -- A Survey

  • 将生成方式从逐词预测拓展为多词、规划生成、潜空间推理等五类新方法
  • 提出系统性分类框架,涵盖连续生成与非Transformer架构等前沿方向
  • 适合关注大模型生成效率与长程规划的研究者与开发者

Next Token Prediction(NTP)范式推动了大语言模型的空前成功,但也导致其长期规划能力差、错误累积和计算效率低等顽疾。随着对替代NTP方法的兴趣日益增长,本综述梳理了新兴的替代生态。我们将这些方法分为五大类:(1) 多词预测,即一次性预测未来多个词;(2) 先规划再生成,提前制定全局高层计划以指导逐词解码;(3) 潜在推理,将自回归过程迁移至连续潜空间;(4) 连续生成方法,用扩散、流匹配或基于能量的方法实现迭代并行优化,取代序列生成;(5) 非Transformer架构,通过固有结构绕过NTP。通过整合这些方法的洞见,本综述提供了一个分类体系,引导研究应对逐词生成的已知局限,推动自然语言处理中新型变革性模型的发展。

原文摘要 · Abstract (English)

The paradigm of Next Token Prediction (NTP) has driven the unprecedented success of Large Language Models (LLMs), but is also the source of their most persistent weaknesses such as poor long-term planning, error accumulation, and computational inefficiency. Acknowledging the growing interest in exploring alternatives to NTP, the survey describes the emerging ecosystem of alternatives to NTP. We categorise these approaches into five main families: (1) Multi-Token Prediction, which targets a block of future tokens instead of a single one; (2) Plan-then-Generate, where a global, high-level plan is created upfront to guide token-level decoding; (3) Latent Reasoning, which shifts the autoregressive process itself into a continuous latent space; (4) Continuous Generation Approaches, which replace sequential generation with iterative, parallel refinement through diffusion, flow matching, or energy-based methods; and (5) Non-Transformer Architectures, which sidestep NTP through their inherent model structure. By synthesizing insights across these methods, this survey offers a taxonomy to guide research into models that address the known limitations of token-level generation to develop new transformative models for natural language processing.

语言模型生成范式大模型架构创新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。