用多个奖励信号迭代微调大模型,提升生成质量
Iterative Foundation Model Fine-Tuning on Multiple Rewards
- 通过多奖励信号迭代优化模型输出
- 在文本、生物序列和分子生成中超越现有方法
- 适合需要多目标优化的生成任务
微调基础模型已成为生成具有特定属性对象的强大方法。强化学习(RL)为此提供了有效框架,使模型能生成最大化给定奖励函数的输出。然而,在文本生成和药物发现等应用中,仅使用单一奖励信号可能不理想,因常需多个评估标准。本文提出一种基于强化学习的多奖励信号微调基础模型的新方法。通过在多个奖励信号间采用迭代微调策略,该方法推广了最先进的基于RL的微调方法。我们进一步提供理论分析,揭示多奖励强化学习微调的表现机制。在文本、生物序列和小分子生成等多个领域进行的实验结果表明,所提算法相比现有最先进基线表现更优。
原文摘要 · Abstract (English)
Fine-tuning foundation models has emerged as a powerful approach for generating objects with specific desired properties. Reinforcement learning (RL) provides an effective framework for this purpose, enabling models to generate outputs that maximize a given reward function. However, in many applications such as text generation and drug discovery, it can be suboptimal to optimize using a single reward signal, as multiple evaluation criteria are often necessary. This paper proposes a novel reinforcement learning-based method for fine-tuning foundation models using multiple reward signals. By employing an iterative fine-tuning strategy across these rewards, our approach generalizes state-of-the-art RL-based methods. We further provide a theoretical analysis that offers insights into the performance of multi-reward RL fine-tuning. Experimental results across diverse domains including text, biological sequence, and small molecule generation, demonstrate the effectiveness of the proposed algorithm compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。