arXiv:2511.13254cs.CL2025-11被引 4

通过加权平均专家模型,用简单算术提升大模型性能

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

  • 按基准测试类别筛选专家模型,非均匀加权融合
  • 在多语言、工具调用等任务上达到顶尖表现
  • 适合追求高效性能提升的模型优化研究者

大语言模型虽能力强大,但训练成本高昂。模型汤(model souping)——即对同架构多个模型权重取平均——是一种无需重新训练即可提升性能的可行方法。本文提出类别专家模型之汤(SoCE),基于基准测试组合识别最优模型候选,采用非均匀加权平均以最大化性能。不同于以往的均匀平均,该方法利用不同评测类别间模型表现的低相关性,为弱相关类别簇分别选出‘专家’模型,并通过优化加权方式融合。实验表明,该方法在多语言、工具调用、数学推理等多个领域均显著提升性能与鲁棒性,在伯克利函数调用排行榜上取得当前最佳结果。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive compute power and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre- and post-training technique that can enhance performance without expensive retraining. In this paper, we introduce Soup Of Category Experts (SoCE), a principled approach for model souping that utilizes benchmark composition to identify optimal model candidates and applies non-uniform weighted averaging to maximize performance. Contrary to previous uniform-averaging approaches, our method leverages the observation that benchmark categories often exhibit low inter-correlations in model performance. SoCE identifies "expert" models for each weakly-correlated category cluster and combines them using optimized weighted averaging rather than uniform weights. We demonstrate that the proposed method improves performance and robustness across multiple domains, including multilingual capabilities, tool calling, and math and achieves state-of-the-art results on the Berkeley Function Calling Leaderboard.

大模型优化模型融合性能提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。