arXiv:2508.03140cs.CLcs.AI2025-08AAAI被引 5

让推理模型与领域模型融合,保持强推理能力同时提升专业表现。

RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior

  • 以推理能力为先验,选择性合并领域知识权重。
  • 在生物医学和金融领域任务上性能提升9.5%和9.2%。
  • 适合需要兼顾推理与专业性的应用开发人员。

具备长链式思维(CoT)能力的大语言模型(即推理模型)通过多步推理展现出卓越的复杂问题求解能力。为在不增加大量计算与数据成本的前提下构建兼具长链式思维能力与领域专业知识的双能力模型,模型融合成为高效方法。然而,现有融合方法在整合领域专用模型与长链式思维模型时,常导致推理能力下降,甚至出现胡言乱语或输出崩溃。为此,本文提出RCP-Merging:一种将长链式思维模型与领域专用模型融合的新框架,该框架将推理模型权重视为基础先验,利用推理能力指标保留核心长链式思维权重,同时选择性融合关键领域专用权重。我们在Qwen2.5-7B、Llama3.1-8B和Qwen2.5-1.5B模型上针对生物医学与金融领域进行了广泛实验。结果表明,RCP-Merging成功实现两类模型融合,在领域任务性能上分别比当前最优方法提升9.5%与9.2%,且未显著损害原始长链式思维推理能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) with long chain-of-thought (CoT) capability, termed Reasoning Models, demonstrate superior intricate problem-solving abilities through multi-step long CoT reasoning. To create a dual-capability model with long CoT capability and domain-specific knowledge without substantial computational and data costs, model merging emerges as a highly resource-efficient method. However, significant challenges lie in merging domain-specific LLMs with long CoT ones since nowadays merging methods suffer from reasoning capability degradation, even gibberish output and output collapse. To overcome this, we introduce RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior, a novel merging framework designed to integrate domain-specific LLMs with long CoT capability, meanwhile maintaining model performance in the original domain. Treating reasoning model weights as foundational prior, our method utilizes a reasoning capability indicator to preserve core long CoT capability model weights while selectively merging essential domain-specific weights. We conducted extensive experiments on Qwen2.5-7B, Llama3.1-8B, and Qwen2.5-1.5B models in BioMedicine and Finance domains. Our results show that RCP-Merging successfully merges a reasoning model with domain-specific ones, improving domain task performance by 9.5% and 9.2% over state-of-the-art methods, without significantly harming the original long CoT reasoning capability.

模型融合推理增强领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。