一种统一算法,让大模型在不同数据下更稳定地优化生成质量。
RL-finetuning LLMs from on- and off-policy data with a single algorithm
- 基于生成一致性设计新算法,兼顾在线与离线数据。
- 在数学推理任务上优于基线,提升模型生成准确性。
- 适合需要高效训练大模型的科研与工程人员。
我们提出一种新型强化学习算法(AGRO,即任意生成奖励优化),用于微调大语言模型。AGRO 基于生成一致性概念——最优策略应满足模型任意生成路径下的内在一致性。通过样本化策略梯度方法推导出求解算法,并提供收敛性理论保证。实验表明,AGRO 在在线与离线设置下均表现有效,在数学推理数据集上的性能超越基线算法。
原文摘要 · Abstract (English)
We introduce a novel reinforcement learning algorithm (AGRO, for Any-Generation Reward Optimization) for fine-tuning large-language models. AGRO leverages the concept of generation consistency, which states that the optimal policy satisfies the notion of consistency across any possible generation of the model. We derive algorithms that find optimal solutions via the sample-based policy gradient and provide theoretical guarantees on their convergence. Our experiments demonstrate the effectiveness of AGRO in both on-policy and off-policy settings, showing improved performance on the mathematical reasoning dataset over baseline algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。