arXiv:2509.24244cs.AI2025-09被引 8

发现大模型合并的缩放规律,可预测加专家带来的性能提升。

Model Merging Scaling Laws in Large Language Models

  • 提出模型大小与专家数量间的幂律关系,揭示合并增益随专家数增加而递减。
  • 实验验证在多种架构和方法下均成立,早期加入专家收益最大,后期增益趋缓。
  • 为模型合并提供可预测的规划工具,适合资源受限下的高效模型设计。

我们研究了基于交叉熵衡量的大语言模型合并的实证缩放规律。尽管合并被广泛使用,但缺乏预测添加专家或扩大模型规模时收益的定量规则。我们发现一个紧凑的幂律关系:模型容量越大,合并的下限越低;而随着专家数量增加,合并效果呈现明显的边际递减。该规律在同域和跨域场景中均成立,能紧密拟合不同架构与方法(如Average、TA、TIES、DARE)的实测曲线,并解释两个稳健现象:大部分增益集中在早期,且随着专家增多,结果方差减小。基于此,我们提出一个简单理论,解释为何增益约按1/k衰减,并将下限与尾部特性关联到基础模型属性及领域多样性。该规律可实现预测性规划:估算达到目标损失所需的专家数、判断停止添加的时机,并在固定预算下权衡扩展基础模型或增加专家——使合并从经验做法转变为计算高效的可规划替代多任务训练方案。这提示分布式生成式AI的一种缩放原则:通过组合专家实现可预测的性能提升,为迈向通用智能系统提供互补路径。

原文摘要 · Abstract (English)

We study empirical scaling laws for language model merging measured by cross-entropy. Despite its wide practical use, merging lacks a quantitative rule that predicts returns as we add experts or scale the model size. We identify a compact power law that links model size and expert number: the size-dependent floor decreases with model capacity, while the merging tail exhibits clear diminishing returns in the number of experts. The law holds in-domain and cross-domain, tightly fits measured curves across diverse architectures and methods (Average, TA, TIES, DARE), and explains two robust regularities: most gains arrive early, and variability shrinks as more experts are included. Building on this, we present a simple theory that explains why gains fall roughly as 1/k and links the floor and tail to properties of the base model and the diversity across domains. This law enables predictive planning: estimate how many experts are needed to reach a target loss, decide when to stop adding experts, and trade off scaling the base model versus adding experts under a fixed budget--turning merging from heuristic practice into a computationally efficient, planable alternative to multitask training. This suggests a scaling principle for distributed generative AI: predictable gains can be achieved by composing specialists, offering a complementary path toward AGI-level systems.

模型合并缩放定律大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。