用基因演化方法自动设计更优的模型架构,提升语言模型性能与效率。
STAR: Synthesis of Tailored Architectures
- 基于线性输入变化系统理论构建可编码的架构搜索空间
- 在多个指标上超越优化过的Transformer和混合模型
- 适合追求高性能、低资源消耗模型的研究者
模型架构的迭代优化是深度学习的核心:Transformer的出现推动了规模扩展,近期模型混合技术进一步拓展了质量与效率的边界。然而,架构优化仍面临挑战且成本高昂。现有自动化或人工方法受限于搜索空间设计不足及生成模式过于简单。本文提出一种新型定制化架构合成方法(STAR),结合线性输入变化系统理论构建新的搜索空间,支持架构基因的分层数值编码。通过无梯度、进化算法对基因进行自动优化与重组,以多目标优化模型质量与效率。利用STAR,我们优化出大规模新架构群体,融合多样计算单元与连接模式,在自回归语言建模任务中,于质量、参数量与推理缓存方面均优于高度优化的Transformer与条纹混合模型。
原文摘要 · Abstract (English)
Iterative improvement of model architectures is fundamental to deep learning: Transformers first enabled scaling, and recent advances in model hybridization have pushed the quality-efficiency frontier. However, optimizing architectures remains challenging and expensive. Current automated or manual approaches fall short, largely due to limited progress in the design of search spaces and due to the simplicity of resulting patterns and heuristics. In this work, we propose a new approach for the synthesis of tailored architectures (STAR). Our approach combines a novel search space based on the theory of linear input-varying systems, supporting a hierarchical numerical encoding into architecture genomes. STAR genomes are automatically refined and recombined with gradient-free, evolutionary algorithms to optimize for multiple model quality and efficiency metrics. Using STAR, we optimize large populations of new architectures, leveraging diverse computational units and interconnection patterns, improving over highly-optimized Transformers and striped hybrid models on the frontier of quality, parameter size, and inference cache for autoregressive language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。