arXiv:2509.19476cs.CL2025-09

首次从行为与内部机制双视角评估模型合并方法。

A Pipeline to Assess Merging Methods via Behavior and Internals

  • 构建新评估流水线,同步分析合并模型的行为与内部表征。
  • 合并后模型在语法形态上表现优于父模型,但性能居中。
  • 行为与内部能力相关性弱,需更全面评估合并效果。

模型合并方法通过融合多个语言模型的权重来提升能力,如领域适配。现有研究多仅从行为角度评估,本文首次提出综合评估框架,同时考察合并模型在下游任务(如MMLU)上的表现及其内部编码的语言能力。我们以Qwen2.5系列中指令微调后的数学与代码适应模型为父模型进行合并,结果表明:合并模型的行为表现通常介于两个父模型之间,但其对形态、句法等语言现象的内部表征能力可能超越父模型。此外,行为评估与内部评估之间的排名相关性较弱。本研究强调,应采用更全面的评估方式,才能真实理解模型合并方法的实际能力与可靠性,避免仅关注表面性能提升。

原文摘要 · Abstract (English)

Merging methods combine the weights of multiple language models (LMs) to leverage their capacities, such as for domain adaptation. While existing studies investigate merged models from a solely behavioral perspective, we offer the first comprehensive view by assessing and connecting their behavior and internals. We present a novel evaluation pipeline that first merges multiple parent LMs, and then evaluates the merged models in comparison to the initial ones based on their behavior on downstream tasks, like MMLU, and the internal encoded linguistic competence. We showcase this pipeline by assessing the merging of instruction fine-tuned with math- and code-adapted LMs from the Qwen2.5 family. Our results show that merging methods impacts behavior and internals differently. While the performance of merged models is typically between that of the two parent models, their encoded information about linguistic phenomena, particularly in morphology and syntax, can surpass the parent models. Moreover, we find weak ranking correlation between this behavior and internal evaluation. With our pipeline and initial results, we emphasize the need for more comprehensive evaluations of model merging methods to gain a faithful understanding of their capabilities and reliability, beyond potential superficial behavioral advances.

模型合并语言模型内部表征评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。