arXiv:2410.18287cs.CLcs.LG2024-10被引 1

从大模型中提取小模型模块,高效定制个性化语言模型。

LEGO: Language Model Building Blocks

  • 通过剪枝技术从大模型提取可复用的小模型组件。
  • 支持任务与用户定制,推理和微调效率高且隐私安全。
  • 结合联邦学习实现模型多样性与数据异构性缓解。

大型语言模型(LLMs)在自然语言处理中至关重要,但其数据收集、预训练、微调和推理成本高昂。特定任务的小型语言模型(SLMs)虽更经济,却缺乏鲁棒性和泛化能力。本文提出LEGO,一种从LLM中提取并重组SLM的新方法。利用先进的LLM剪枝策略,可生成针对任务和用户定制的高效SLM构建块,兼具低开销与高隐私保护。LEGO结合联邦学习与新型聚合方案,重建大模型时保持鲁棒性,同时避免高成本。实验验证了LEGO的多功能性,能支持模型异构性,缓解数据异构性影响,且维持原有大模型的性能稳健性。

原文摘要 · Abstract (English)

Large language models (LLMs) are essential in natural language processing (NLP) but are costly in data collection, pre-training, fine-tuning, and inference. Task-specific small language models (SLMs) offer a cheaper alternative but lack robustness and generalization. This paper proposes LEGO, a novel technique to extract SLMs from an LLM and recombine them. Using state-of-the-art LLM pruning strategies, we can create task- and user-specific SLM building blocks that are efficient for fine-tuning and inference while also preserving user data privacy. LEGO utilizes Federated Learning and a novel aggregation scheme for the LLM reconstruction, maintaining robustness without high costs and preserving user data privacy. We experimentally demonstrate the versatility of LEGO, showing its ability to enable model heterogeneity and mitigate the effects of data heterogeneity while maintaining LLM robustness.

语言模型模型剪枝联邦学习小型化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。