arXiv:2511.11572cs.GLcs.CL2025-11被引 1

梳理大模型架构与规模经济,揭示算力、内存与成本关系

LLM Architecture, Scaling Laws, and Economics: A Quick Summary

  • 基于标准Transformer结构总结主流LLM的QKV自注意力设计
  • 给出2025年不同规模模型的算力与内存成本估算数据
  • 适合关注大模型部署成本与技术演进路线的研究者参考

本文简要总结了当前大型语言模型(LLMs)的标准架构,包括典型Transformer的结构。给出了算力(浮点运算次数,flops)和内存(参数加数据)的缩放定律,并提供了2025年各类规模现役大模型参数成本的粗略估算。同时讨论了DeepSeek是否应被视为特例。内容并无新意,但此类信息在现有文献中尚未以综述形式集中呈现。

原文摘要 · Abstract (English)

The current standard architecture of Large Language Models (LLMs) with QKV self-attention is briefly summarized, including the architecture of a typical Transformer. Scaling laws for compute (flops) and memory (parameters plus data) are given, along with their present (2025) rough cost estimates for the parameters of present LLMs of various scales, including discussion of whether DeepSeek should be viewed as a special case. Nothing here is new, but this material seems not otherwise readily available in summary form.

大模型架构缩放定律成本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。