不受限的模型融合让多个专家大模型联手提升推理能力。
Unconstrained Model Merging for Enhanced LLM Reasoning
- 突破架构限制,统一融合同质与异质模型的权重
- 在9个推理优化模型上实现组合式推理超越简单叠加
- 适合想低成本构建强推理模型的研究者与开发者
近期在构建领域专用大语言模型(LLMs)方面取得显著进展,尤其在逻辑推理和多步问题求解等需要推理能力的任务中表现突出。然而,由于需依赖专有数据和大量计算资源,打造一个全能型大模型仍具挑战性。作为更节省资源的替代方案,我们探索将多个专家模型合并为单一模型的可能性。现有研究主要集中在通用大模型或相同架构与规模的模型上,而本文提出一种无约束模型融合框架,可兼容同质与异质模型架构,聚焦于推理任务。针对同质模型设计细粒度逐层权重融合策略,异质模型则基于指令-响应微调数据中的概率分布知识进行融合。在7个基准测试和9个推理优化的LLM上,发现组合推理能力通过融合涌现,远超简单相加效果。我们认为无约束模型融合可成为去中心化大模型的基础,标志着从现有集中式大模型框架的重要演进,有望提升参与广度并推动人工智能领域进一步发展。
原文摘要 · Abstract (English)
Recent advancements in building domain-specific large language models (LLMs) have shown remarkable success, especially in tasks requiring reasoning abilities like logical inference over complex relationships and multi-step problem solving. However, creating a powerful all-in-one LLM remains challenging due to the need for proprietary data and vast computational resources. As a resource-friendly alternative, we explore the potential of merging multiple expert models into a single LLM. Existing studies on model merging mainly focus on generalist LLMs instead of domain experts, or the LLMs under the same architecture and size. In this work, we propose an unconstrained model merging framework that accommodates both homogeneous and heterogeneous model architectures with a focus on reasoning tasks. A fine-grained layer-wise weight merging strategy is designed for homogeneous models merging, while heterogeneous model merging is built upon the probabilistic distribution knowledge derived from instruction-response fine-tuning data. Across 7 benchmarks and 9 reasoning-optimized LLMs, we reveal key findings that combinatorial reasoning emerges from merging which surpasses simple additive effects. We propose that unconstrained model merging could serve as a foundation for decentralized LLMs, marking a notable progression from the existing centralized LLM framework. This evolution could enhance wider participation and stimulate additional advancement in the field of artificial intelligence, effectively addressing the constraints posed by centralized models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。