提出可弹性搜索的紧凑语言模型架构,提升效率与性能。
Elastic Architecture Search for Efficient Language Models
- 设计灵活搜索空间,支持动态调整维度和注意力头数。
- 在掩码语言建模与因果语言建模任务中超越现有方法。
- 适合追求高效推理与低资源部署的语言模型研究者。
随着大规模预训练语言模型在自然语言理解任务中的重要性日益增加,其巨大的计算与内存开销带来了显著的经济与环境负担。本文提出弹性语言模型(ELM),一种面向紧凑语言模型的新型神经架构搜索(NAS)方法。ELM通过引入高效的Transformer模块和可动态调整维度与注意力头数的模块,扩展了现有NAS的搜索空间,提升了搜索过程的效率与灵活性,从而实现更全面有效的架构探索。此外,我们设计了新的知识蒸馏损失函数,以保留各模块的独特特征,增强架构选择过程中的判别能力。在掩码语言建模与因果语言建模任务上的实验表明,由ELM发现的模型显著优于现有方法。
原文摘要 · Abstract (English)
As large pre-trained language models become increasingly critical to natural language understanding (NLU) tasks, their substantial computational and memory requirements have raised significant economic and environmental concerns. Addressing these challenges, this paper introduces the Elastic Language Model (ELM), a novel neural architecture search (NAS) method optimized for compact language models. ELM extends existing NAS approaches by introducing a flexible search space with efficient transformer blocks and dynamic modules for dimension and head number adjustment. These innovations enhance the efficiency and flexibility of the search process, which facilitates more thorough and effective exploration of model architectures. We also introduce novel knowledge distillation losses that preserve the unique characteristics of each block, in order to improve the discrimination between architectural choices during the search process. Experiments on masked language modeling and causal language modeling tasks demonstrate that models discovered by ELM significantly outperform existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。