Bone Soup通过融合多目标模型实现可控生成,适应用户多样化需求。
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
- 基于多目标强化学习训练多个主干模型,融合不同目标奖励信号。
- 在测试时根据用户偏好,用对称循环矩阵生成混合系数,实现精准控制。
- 相比传统方法更优的帕累托最优性,适合需要灵活响应的生成任务。
用户需求高度多样,当前研究面临如何在测试时快速适应不同需求的同时实现可控多目标生成的挑战。现有方法如Rewarded Soup仅对单目标微调的语言模型进行合并,忽视了不同目标间的相互影响,导致性能受限。为此,本文提出Bone Soup,一种新型模型合并方法:首先通过多目标强化学习训练多个主干模型,每个模型由一组基向量形式的主干奖励信号引导;这些奖励通过规则化构造方式生成,确保模型位于帕累托前沿。随后,Bone Soup利用对称循环矩阵映射生成合并系数,按用户偏好融合主干模型。大量实验表明,Bone Soup在可控多目标生成中展现出优异的可控性与帕累托最优性,为测试阶段应对多样化用户需求提供了更高效、有效的新途径。
原文摘要 · Abstract (English)
User information needs are often highly diverse and varied. A key challenge in current research is how to achieve controllable multi-objective generation while enabling rapid adaptation to accommodate diverse user demands during test time. Existing solutions, such as Rewarded Soup, focus on merging language models individually tuned on single objectives. While easy to implement and widely used, these approaches face limitations in achieving optimal performance due to their disregard for the impacts of competing objectives on model tuning. To address this issue, we propose Bone Soup, a novel model merging approach that first seeks a series of backbone models by considering the impacts of multiple objectives and then makes the soup (i.e., merge the backbone models). Specifically, Bone Soup begins by training multiple backbone models for different objectives using multi-objective reinforcement learning. Each backbone model is guided by a combination of backbone reward signals. To ensure that these models are optimal for the Pareto front, the backbone rewards are crafted by combining standard reward functions into basis vectors, which can then be modified through a rule-based construction method. Bone Soup leverages a symmetric circulant matrix mapping to generate the merging coefficients, which are used to merge the backbone models according to user preferences. Extensive experimental results demonstrate that Bone Soup exhibits strong controllability and Pareto optimality in controllable multi-objective generation, providing a more effective and efficient approach to addressing diverse user needs at test time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。