arXiv:2508.01148cs.LG2025-08被引 2

用知识蒸馏预处理任务向量,提升模型融合的鲁棒性

DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging

  • 通过知识蒸馏调整任务向量幅度和源模型置信度
  • 在任务向量差异大或模型置信度低时仍能显著提效
  • 适合需融合多不靠谱模型的实用场景

模型融合已成为一种高效灵活的多任务学习范式,近年来涌现多种方法。然而,这些先进方法通常在对融合友好的基准测试中评估,其在更真实场景下的鲁棒性仍不明朗。本文首次分析模型融合方法的脆弱性,发现两个关键问题:(1) 任务向量模长差异大,(2) 源模型置信度低。为此,提出DisTaC(Distillation for Task vector Conditioning),一种在融合前预调节任务向量的新方法。DisTaC利用知识蒸馏调整任务向量模长并提升源模型置信度,同时保留其核心任务知识。大量实验表明,经DisTaC预处理后,先进融合技术可成功整合原本会失败的模型,实现显著性能提升。

原文摘要 · Abstract (English)

Model merging has emerged as an efficient and flexible paradigm for multi-task learning, with numerous methods being proposed in recent years. However, these state-of-the-art techniques are typically evaluated on benchmark suites that are highly favorable to model merging, and their robustness in more realistic settings remains largely unexplored. In this work, we first investigate the vulnerabilities of model-merging methods and pinpoint the source-model characteristics that critically underlie them. Specifically, we identify two factors that are particularly harmful to the merging process: (1) disparities in task vector norms, and (2) the low confidence of the source models. To address this issue, we propose DisTaC (Distillation for Task vector Conditioning), a novel method that pre-conditions these problematic task vectors before the merge. DisTaC leverages knowledge distillation to adjust a task vector's norm and increase source-model confidence while preserving its essential task-specific knowledge. Our extensive experiments demonstrate that by pre-conditioning task vectors with DisTaC, state-of-the-art merging techniques can successfully integrate models exhibiting the harmful traits -- where they would otherwise fail -- achieving significant performance gains.

模型融合知识蒸馏多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。