arXiv:2601.02256cs.CVcs.LG2026-01被引 2

解决视觉自回归生成中异步策略冲突问题,提升强化学习训练稳定性。

VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation

  • 引入中间奖励与动态时间权重,缓解生成过程中的策略冲突。
  • 在多个数据集上显著提升样本质量与目标对齐度,优于基线方法。
  • 适合研究视觉生成、强化学习与复杂序列建模的科研人员。

视觉生成主要分为自回归(AR)、扩散和视觉自回归(VAR)三种范式。与AR和扩散模型不同,VAR在生成各步骤中处理异构输入结构,导致严重的异步策略冲突。这一问题在强化学习(RL)场景下尤为突出,引发训练不稳定与对齐效果不佳。为此,我们提出一种新框架,通过显式管理这些冲突来增强组相对策略优化(GRPO)。方法包含三个协同组件:1)稳定早期生成的中间奖励;2)动态时间步重加权方案,实现精确信用分配;3)基于奖励反馈学习(ReFL)原理设计的新型掩码传播算法,可实现时空维度上的优化效应隔离。实验表明,该方法在样本质量和目标对齐方面显著优于原始GRPO基线,使VAR模型的优化更加鲁棒有效。

原文摘要 · Abstract (English)

Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous input structures across their generation steps, which creates severe asynchronous policy conflicts. This issue becomes particularly acute in reinforcement learning (RL) scenarios, leading to unstable training and suboptimal alignment. To resolve this, we propose a novel framework to enhance Group Relative Policy Optimization (GRPO) by explicitly managing these conflicts. Our method integrates three synergistic components: 1) a stabilizing intermediate reward to guide early-stage generation; 2) a dynamic time-step reweighting scheme for precise credit assignment; and 3) a novel mask propagation algorithm, derived from principles of Reward Feedback Learning (ReFL), designed to isolate optimization effects both spatially and temporally. Our approach demonstrates significant improvements in sample quality and objective alignment over the vanilla GRPO baseline, enabling robust and effective optimization for VAR models.

视觉生成强化学习序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。