分析大模型架构趋同现象及其在不同设置下的表现差异。
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
- 追踪操作演进路径,梳理大模型架构发展脉络。
- 同一模型在不同超参数或部署环境下发行为迥异。
- 基于RTX 6000评测显示,硬件与配置显著影响性能。
注意力机制与Transformer架构的出现,使上下文自然的文本生成成为可能,并将处理完整源信息的负担压缩为单一向量。基于这两项核心思想,模型规模持续扩大以容纳更精确、全面的信息,当前最先进大模型参数量已达约700亿。随着模型规模增长,对存储和计算能力的需求急剧上升,推动了高带宽内存与加速器的发展,以及多种旨在满足需求的模型架构设计。我们观察到大模型架构正趋于收敛。本文通过追溯操作改进的历史,分析这些收敛架构在层配置、操作机制及模型规模上的表现,考虑多种超参数设置。同时,利用搭载最新Ada Lovelace架构的RTX 6000,总结大模型在不同超参数设置下的性能趋势。结论表明,即使相同模型,其行为也会因超参数或部署于服务器/边缘环境而显著不同。
原文摘要 · Abstract (English)
The advent of the Attention mechanism and Transformer architecture enables contextually natural text generation and compresses the burden of processing entire source information into singular vectors. Based on these two main ideas, model sizes gradually increases to accommodate more precise and comprehensive information, leading to the current state-of-the-art LLMs being very large, with parameters around 70 billion. As the model sizes are growing, the demand for substantial storage and computational capacity increases. This leads to the development of high-bandwidth memory and accelerators, as well as a variety of model architectures designed to meet these requirements. We note that LLM architectures have increasingly converged. This paper analyzes how these converged architectures perform in terms of layer configurations, operational mechanisms, and model sizes, considering various hyperparameter settings. In this paper, we conduct a concise survey of the history of LLMs by tracing the evolution of their operational improvements. Furthermore, we summarize the performance trends of LLMs under various hyperparameter settings using the RTX 6000, which features the state-of-the-art Ada Lovelace architecture. We conclude that even the same model can exhibit different behaviors depending on the hyperparameters or whether it is deployed in server or edge environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。