无需训练和数据,通过语义对齐融合语言模型
SeMe: Training-Free Language Model Merging via Semantic Alignment
- 基于潜在语义对齐,在层级别实现细粒度模型融合
- 在多个架构和任务上表现优于现有方法,且不依赖外部数据
- 适合需要高效组合大模型能力的研究者与开发者
尽管语言模型在多种任务中表现出色,但没有单一模型始终领先,因此需高效方法融合其优势而无需昂贵的重新训练。现有模型融合技术如参数平均和任务引导融合常依赖数据计算或无法保留内部知识,限制了鲁棒性与可扩展性。我们提出SeMe(基于语义的融合),一种全新的、无需数据且免训练的方法,通过潜在语义对齐在细粒度、逐层层面融合语言模型。与以往工作不同,SeMe不仅保留模型行为,还显式稳定内部知识,填补了语言模型融合的关键空白。在多种架构和任务上的广泛实验表明,SeMe在性能与效率上均优于现有方法,同时消除对外部数据的依赖。本工作建立了一种新的知识感知模型融合范式,并为理解语言模型的语义结构提供了洞见,推动更可扩展、可解释的模型组合发展。
原文摘要 · Abstract (English)
Despite the remarkable capabilities of Language Models (LMs) across diverse tasks, no single model consistently outperforms others, necessitating efficient methods to combine their strengths without expensive retraining. Existing model merging techniques, such as parameter averaging and task-guided fusion, often rely on data-dependent computations or fail to preserve internal knowledge, limiting their robustness and scalability. We introduce SeMe (Semantic-based Merging), a novel, data-free, and training-free approach that leverages latent semantic alignment to merge LMs at a fine-grained, layer-wise level. Unlike prior work, SeMe not only preserves model behaviors but also explicitly stabilizes internal knowledge, addressing a critical gap in LM fusion. Through extensive experiments across diverse architectures and tasks, we demonstrate that SeMe outperforms existing methods in both performance and efficiency while eliminating reliance on external data. Our work establishes a new paradigm for knowledge-aware model merging and provides insights into the semantic structure of LMs, paving the way for more scalable and interpretable model composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。