弱模型提背景,强模型做推理,协同提升专业任务表现
Synergistic Weak-Strong Collaboration by Aligning Preferences
- 弱模型提供领域信息,强模型基于其优化推理
- 协同效果显著优于单个模型,提升关键任务性能
- 通过偏好对齐微调弱模型,增强协作效率
当前大语言模型在通用推理上表现优异,但在需要专有或领域知识的特定任务中表现不佳。为解决此问题,我们提出一种弱-强模型协同框架:弱模型针对特定领域生成初始草稿与背景信息,强模型利用其先进推理能力进行优化,从而拓展大模型在关键专业任务中的应用。为优化协作,我们引入协同反馈机制,量化弱模型贡献并构建偏好对,指导弱模型的偏好微调。在三个领域的实验验证表明,该框架显著优于单一模型表现;进一步对齐弱模型与协同偏好,可进一步提升整体性能。
原文摘要 · Abstract (English)
Current Large Language Models (LLMs) excel in general reasoning yet struggle with specialized tasks requiring proprietary or domain-specific knowledge. Fine-tuning large models for every niche application is often infeasible due to black-box constraints and high computational overhead. To address this, we propose a collaborative framework that pairs a specialized weak model with a general strong model. The weak model, tailored to specific domains, produces initial drafts and background information, while the strong model leverages its advanced reasoning to refine these drafts, extending LLMs' capabilities to critical yet specialized tasks. To optimize this collaboration, we introduce a collaborative feedback to fine-tunes the weak model, which quantifies the influence of the weak model's contributions in the collaboration procedure and establishes preference pairs to guide preference tuning of the weak model. We validate our framework through experiments on three domains. We find that the collaboration significantly outperforms each model alone by leveraging complementary strengths. Moreover, aligning the weak model with the collaborative preference further enhances overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。