arXiv:2605.15803cs.CVcs.LG2026-05中稿 · ICML被引 3

通过嵌入扰动保持生成模型优化中的多样性信号

Embedding-perturbed Exploration Preference Optimization for Flow Models

论文配图:Embedding-perturbed Exploration Preference Optimization for Flow Models
图 1 · 摘自论文原文
  • 在样本组内引入嵌入层扰动,维持优化所需差异性
  • 实验显示显著优于现有基线,更贴近人类偏好
  • 适合需要稳定训练的生成模型对齐任务

近期进展已确立强化学习(RL)作为对齐生成模型与人类意图的关键范式。然而,基于群体的优化框架(如GRPO)面临关键局限:组内方差迅速衰减。随着组内样本差异性消失,方差趋近于零,导致优化所需的梯度信号丧失,引发训练不稳定或策略过早停滞、奖励欺骗。现有方法如改变初始噪声或增大群体规模,常无法根本解决此问题,造成训练不稳或收益递减。为此,我们提出嵌入扰动探索偏好优化(E²PO),一种通过嵌入级扰动维持优化过程稳定性的新框架。该方法在样本组内引入结构化嵌入级扰动,确保鲁棒方差持续存在,从而保留判别性信号。大量实验表明,本方法显著优于当前最优基线,在更忠实对齐人类偏好方面表现优异。

原文摘要 · Abstract (English)

Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-based optimization frameworks (e.g., GRPO) face a critical limitation: the rapid decay of intra-group variance. As the distinctiveness among samples within a group diminishes, the variance approaches zero. This eliminates the very learning signal required for optimization, rendering the process unstable and forcing the policy into premature stagnation or reward hacking. Existing strategies, such as varying the initial noise or increasing group sizes, often fail to address this fundamental issue, resulting in training instability or diminishing returns. To overcome these challenges, we propose $\textbf{Embedding-perturbed Exploration Preference Optimization (}E^2\textbf{PO)}$, a novel framework that sustains optimization through embedding-level perturbation. Our method introduces structured, embedding-level perturbations within sample groups, guaranteeing a robust variance that preserves the discriminative signal throughout the training process. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art baselines, achieving a more faithful alignment with human preference.

生成模型强化学习偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。