提出统一框架GRAO,融合监督与强化学习优势提升模型对齐效率。
Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- 采用多样本生成和组内相对优势加权,实现更精准的奖励评估。
- 在多个任务上相比SFT、DPO等基线提升超5%,最高达57.7%。
- 适合追求高效对齐优化的模型训练者,尤其关注样本效率的场景。
对齐方法已成为提升语言模型对齐能力的关键路径。尽管监督微调(SFT)通过直接的词元级损失干预加速收敛,但受限于离线策略轨迹;而强化学习(RL)虽能探索优化策略,却存在样本效率低、依赖高质量基础模型的问题。为此,我们提出统一框架GRAO(Group Relative Alignment Optimization),通过三项创新融合SFT与RL的优势:1)多样本生成策略,借助奖励反馈实现对比质量评估;2)新型组内直接对齐损失,利用组内相对优势加权;3)基于成对偏好动态的参考感知参数更新。理论分析证明了GRAO的收敛性及样本效率优势。在复杂人类对齐任务上的全面评估显示,GRAO相比SFT、DPO、PPO和GRPO基线分别实现57.70%、17.65%、7.95%和5.18%的相对提升。本工作为语言模型的高效能力演化提供了理论依据与实证支持。
原文摘要 · Abstract (English)
Alignment methodologies have emerged as a critical pathway for enhancing language model alignment capabilities. While SFT (supervised fine-tuning) accelerates convergence through direct token-level loss intervention, its efficacy is constrained by offline policy trajectory. In contrast, RL(reinforcement learning) facilitates exploratory policy optimization, but suffers from low sample efficiency and stringent dependency on high-quality base models. To address these dual challenges, we propose GRAO (Group Relative Alignment Optimization), a unified framework that synergizes the respective strengths of SFT and RL through three key innovations: 1) A multi-sample generation strategy enabling comparative quality assessment via reward feedback; 2) A novel Group Direct Alignment Loss formulation leveraging intra-group relative advantage weighting; 3) Reference-aware parameter updates guided by pairwise preference dynamics. Our theoretical analysis establishes GRAO's convergence guarantees and sample efficiency advantages over conventional approaches. Comprehensive evaluations across complex human alignment tasks demonstrate GRAO's superior performance, achieving 57.70\%,17.65\% 7.95\% and 5.18\% relative improvements over SFT, DPO, PPO and GRPO baselines respectively. This work provides both a theoretically grounded alignment framework and empirical evidence for efficient capability evolution in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。