arXiv:2507.09955cs.AI2025-07被引 34

DeepSeek开源大模型以低成本高效率推动AI范式变革

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models

  • 提出MLA、MoE等新算法,优化模型架构与训练效率
  • 实现高性能推理与系统级优化,降低部署成本
  • 适合关注国产大模型技术突破的研究者与开发者

DeepSeek是一家中国人工智能初创公司,其发布的V3和R1系列模型因低成本、高性能及开源优势引发全球关注。本文回顾大模型发展中的范式演进,涵盖主流大语言模型(LLM)范式与DeepSeek范式。重点介绍DeepSeek提出的创新算法,包括多头潜在注意力(MLA)、专家混合(MoE)、多标记预测(MTP)和组相对策略优化(GRPO)。文章进一步探讨其在大模型扩展、训练、推理及系统级优化架构方面的工程突破。同时分析了DeepSeek模型对当前竞争性AI格局的影响,与主流大语言模型在多个领域进行对比。最后,总结DeepSeek创新带来的启示,并展望未来大模型在数据、训练与推理方面的技术与工程发展趋势。

原文摘要 · Abstract (English)

DeepSeek, a Chinese Artificial Intelligence (AI) startup, has released their V3 and R1 series models, which attracted global attention due to their low cost, high performance, and open-source advantages. This paper begins by reviewing the evolution of large AI models focusing on paradigm shifts, the mainstream Large Language Model (LLM) paradigm, and the DeepSeek paradigm. Subsequently, the paper highlights novel algorithms introduced by DeepSeek, including Multi-head Latent Attention (MLA), Mixture-of-Experts (MoE), Multi-Token Prediction (MTP), and Group Relative Policy Optimization (GRPO). The paper then explores DeepSeek engineering breakthroughs in LLM scaling, training, inference, and system-level optimization architecture. Moreover, the impact of DeepSeek models on the competitive AI landscape is analyzed, comparing them to mainstream LLMs across various fields. Finally, the paper reflects on the insights gained from DeepSeek innovations and discusses future trends in the technical and engineering development of large AI models, particularly in data, training, and reasoning.

大模型开源算法创新工程优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。