提出结构化方法,精准分配专家模型合并时的容量资源。
Signature-Guided Capacity Occupancy for Dense Expert Merging

- 基于冲突签名确定各层可分配容量。
- 根据领域需求分配容量,平均提升15.0%性能。
- 无需调参搜索,适合高效集成多领域模型。
密集专家合并将领域专用语言模型整合为单一检查点,通常通过在权重空间中引入任务向量支持实现。然而,现有方法仅部分解决了三个关键问题:从何处开启跨专家冲突中的层容量、由谁占据该容量以满足领域需求、以及如何不依赖高成本调优方案来接纳结果支持。为此,我们提出SigMerge(签名引导容量占用)框架,用于密集专家合并的结构化容量分配。从一个密集基础合并开始,冲突签名确定每层容量,正向基合并偏差设定各领域的容量份额,顺序占用规则则依据层-领域预算逐个接纳专家增量。在21组配对实验中(涵盖七个密集基础合并与三个模型池),SigMerge在所有设置中均表现更优(平均提升15.0%),且在六种合并方法中取得最佳平均排名(1.67),超越三类合并基线。
原文摘要 · Abstract (English)
Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space. However, this admission is governed by three decisions that existing methods answer only partially: where to open layer capacity from cross-expert conflict, who should occupy that capacity based on domain demand, and how to admit the resulting support without relying on costly recipe search. To tackle these issues, we propose SigMerge (Signature-Guided Capacity Occupancy), a structured capacity assignment framework for dense expert merging. Starting from a dense base merge, conflict signatures set each layer's capacity from cross-expert conflict, positive base-merge deficits set each domain's share of that capacity, and a sequential occupancy rule admits each expert delta up to the resulting layer-domain budget. Across 21 paired settings spanning seven dense base merges and three model pools, SigMerge improves every one (by 15.0% on average) and achieves the best average rank (1.67) among six merging methods, outperforming three categories of merging baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。