arXiv:2502.10436cs.NEcs.AI2025-02ICML被引 11

让普通显卡也能高效实现多任务模型融合,性能不降反升。

MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs

  • 用简化数据集+项目反应理论估算模型能力,降低评估开销50倍
  • 在单张消费级显卡上实现顶尖跨语言模型融合效果
  • 适合想低成本实验多任务模型融合的研究者和开发者

进化式模型融合能构建高性能多任务模型,但对消费级硬件仍过于昂贵。我们提出MERGE³,一个高效框架,使单张GPU即可实现进化融合,将适应度计算成本降低50倍,同时保持性能。该方法通过提取精简评估数据集、利用项目反应理论(IRT)估计模型能力,并基于IRT的性能预测器演化最优融合方案。我们的方法在多语言与跨语言融合任务中达到当前最优表现,显著降低计算开销。提供理论保障与开源库,推动高质量模型融合的普及。

原文摘要 · Abstract (English)

Evolutionary model merging enables the creation of high-performing multi-task models but remains computationally prohibitive for consumer hardware. We introduce MERGE$^3$, an efficient framework that makes evolutionary merging feasible on a single GPU by reducing fitness computation costs 50$\times$ while preserving performance. MERGE$^3$ achieves this by Extracting a reduced dataset for evaluation, Estimating model abilities using Item Response Theory (IRT), and Evolving optimal merges via IRT-based performance estimators. Our method enables state-of-the-art multilingual and cross-lingual merging, transferring knowledge across languages with significantly lower computational overhead. We provide theoretical guarantees and an open-source library, democratizing high-quality model merging.

模型融合进化计算低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。