arXiv:2504.07158cs.LGcs.CL2025-04被引 1

27.5亿激活参数模型实现强推理与全面能力兼顾

Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models

  • 通过高质量数据与创新训练方法,压缩MoE模型但保留强推理力
  • 推理能力媲美70亿参数模型,通用能力更优
  • 适合需要轻量高效且多任务兼容的部署场景

本技术报告介绍Ring-Lite-Distill,一个基于开源Mixture-of-Experts(MoE)大语言模型Ling-Lite的轻量级推理模型。研究表明,通过精心筛选高质量数据并采用巧妙训练范式,可在仅27.5亿激活参数下,使紧凑的MoE模型获得卓越推理能力,建立高效轻量推理架构。该模型不仅在高难度数学问题求解等高级推理任务中表现突出,更实现了对不同难度推理任务的广泛覆盖,同时保持指令遵循、工具使用和知识保留等通用能力。实验表明,Ring-Lite-Distill的推理能力达到DeepSeek-R1-Distill-Qwen-7B水平,而通用能力显著优于后者。模型已在Hugging Face开放获取。

原文摘要 · Abstract (English)

This technical report presents Ring-Lite-Distill, a lightweight reasoning model derived from our open-source Mixture-of-Experts (MoE) Large Language Models (LLMs) Ling-Lite. This study demonstrates that through meticulous high-quality data curation and ingenious training paradigms, the compact MoE model Ling-Lite can be further trained to achieve exceptional reasoning capabilities, while maintaining its parameter-efficient architecture with only 2.75 billion activated parameters, establishing an efficient lightweight reasoning architecture. In particular, in constructing this model, we have not merely focused on enhancing advanced reasoning capabilities, exemplified by high-difficulty mathematical problem solving, but rather aimed to develop a reasoning model with more comprehensive competency coverage. Our approach ensures coverage across reasoning tasks of varying difficulty levels while preserving generic capabilities, such as instruction following, tool use, and knowledge retention. We show that, Ring-Lite-Distill's reasoning ability reaches a level comparable to DeepSeek-R1-Distill-Qwen-7B, while its general capabilities significantly surpass those of DeepSeek-R1-Distill-Qwen-7B. The models are accessible at https://huggingface.co/inclusionAI

轻量化模型推理能力MoE架构参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。