arXiv:2502.02421cs.CLcs.AI2025-02NeurIPS被引 24

通过激活信息优化大模型合并,性能最高提升40%

Activation-Informed Merging of Large Language Models

  • 融合模型激活空间信息,智能筛选关键权重
  • 在多个基准上实现性能提升,最高达40%
  • 适配任意现有合并方法,适合追求高效模型集成的研究者

模型合并是一种将多个微调过的大型语言模型(LLMs)的参数与嵌入进行整合的方法,能在保持计算效率的同时提升跨任务表现。本文提出激活信息引导的合并(AIM),将LLM激活空间的信息融入合并过程,以增强性能与鲁棒性。AIM作为一种灵活且可互补的方案,适用于任何现有合并方法,其设计基于持续学习(CL)和模型压缩原理,旨在保留基础模型的关键权重。通过使用任务无关的校准集,AIM在合并过程中有选择地优先保留重要权重。实验表明,AIM显著提升了合并模型在多个基准上的表现,结果表明考虑激活空间信息可使模型合并策略取得实质性进展,性能最高提升达40%。

原文摘要 · Abstract (English)

Model merging, a method that combines the parameters and embeddings of multiple fine-tuned large language models (LLMs), offers a promising approach to enhance model performance across various tasks while maintaining computational efficiency. This paper introduces Activation-Informed Merging (AIM), a technique that integrates the information from the activation space of LLMs into the merging process to improve performance and robustness. AIM is designed as a flexible, complementary solution that is applicable to any existing merging method. It aims to preserve critical weights from the base model, drawing on principles from continual learning (CL) and model compression. Utilizing a task-agnostic calibration set, AIM selectively prioritizes essential weights during merging. We empirically demonstrate that AIM significantly enhances the performance of merged models across multiple benchmarks. Our findings suggest that considering the activation-space information can provide substantial advancements in the model merging strategies for LLMs, with up to a 40% increase in benchmark performance.

大模型合并激活信息性能提升模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。