arXiv:2607.16051cs.CLcs.AI2026-07被引 1

Loops Transformer 20倍计算量,性能超越传统大模型

Loop the Loopies!

论文配图:Loop the Loopies!
图 1 · 摘自论文原文
  • 用专家混合架构实现高效循环推理,仅激活部分参数
  • 200亿参数模型在同等算力下超越300亿参数基线
  • 后训练方法提升推理能力,达前沿水平

我们提出Loopie系列,包含两个基于专家混合(MoE)的模型:一个200亿参数、激活参数20亿的模型,以及一个60亿参数、激活参数6亿的模型。长期以来,循环变压器(Looped Transformers)面临挑战:当预训练算力增加N倍时,参数量增加N倍通常比将模型循环N次表现更好。Loopie解决了这一难题。大量消融实验(包括与300亿参数-30亿激活参数模型的对比)表明,Loopie在相同算力预算下显著优于普通Transformer基线。通过一种新颖的后训练方法,Loopie展现出强大的推理能力,达到前沿水平的推理性能。

原文摘要 · Abstract (English)

We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.

MoE推理能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。