arXiv:2512.24618cs.CL2025-12被引 6

1.96B小模型具备强推理与规划能力,适合长任务智能体。

Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models

  • 从零训练构建紧凑架构,支持128k上下文,内存占用低
  • 用11万亿词分阶段训练,逐步提升科学与智能体能力
  • 在数学、编程等任务中内化计划与反思行为,性能超同类小模型

我们提出 Youtu-LLM,一个仅1.96B参数的轻量级大语言模型,兼具高计算效率与原生智能体能力。不同于依赖蒸馏的小模型,Youtu-LLM 从头预训练,系统性培养推理与规划能力。关键技术包括:(1) 基于密集多隐式注意力(MLA)架构,采用面向STEM的词汇表,支持128k上下文窗口,实现低内存下的长程推理与状态追踪;(2) 设计“常识-STEM-智能体”渐进式训练课程,基于约11万亿词数据集,分阶段调整预训练数据分布,确保模型获得深层认知能力而非表面对齐;(3) 在中段训练中引入多样化数据构造策略,合成数学、编程、工具使用等领域的丰富任务轨迹,使模型有效内化规划与反思行为。大量评估表明,Youtu-LLM 在小于20亿参数的模型中达到新SOTA,通用基准表现媲美更大模型,智能体任务显著超越现有最优基线,证明轻量模型可具备强大内在智能体潜力。

原文摘要 · Abstract (English)

We introduce Youtu-LLM, a lightweight yet powerful language model that harmonizes high computational efficiency with native agentic intelligence. Unlike typical small models that rely on distillation, Youtu-LLM (1.96B) is pre-trained from scratch to systematically cultivate reasoning and planning capabilities. The key technical advancements are as follows: (1) Compact Architecture with Long-Context Support: Built on a dense Multi-Latent Attention (MLA) architecture with a novel STEM-oriented vocabulary, Youtu-LLM supports a 128k context window. This design enables robust long-context reasoning and state tracking within a minimal memory footprint, making it ideal for long-horizon agent and reasoning tasks. (2) Principled "Commonsense-STEM-Agent" Curriculum: We curated a massive corpus of approximately 11T tokens and implemented a multi-stage training strategy. By progressively shifting the pre-training data distribution from general commonsense to complex STEM and agentic tasks, we ensure the model acquires deep cognitive abilities rather than superficial alignment. (3) Scalable Agentic Mid-training: Specifically for the agentic mid-training, we employ diverse data construction schemes to synthesize rich and varied trajectories across math, coding, and tool-use domains. This high-quality data enables the model to internalize planning and reflection behaviors effectively. Extensive evaluations show that Youtu-LLM sets a new state-of-the-art for sub-2B LLMs. On general benchmarks, it achieves competitive performance against larger models, while on agent-specific tasks, it significantly surpasses existing SOTA baselines, demonstrating that lightweight models can possess strong intrinsic agentic capabilities.

轻量模型智能体长上下文推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。