arXiv:2505.23331cs.CVcs.AI2025-05被引 5

用强化学习提升视觉自回归模型生成质量与风格控制能力

Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization

  • 采用组相对策略优化方法,基于美学评分和CLIP嵌入进行微调
  • 生成图像质量显著提升,可精准控制风格且突破ImageNet分布限制
  • 适合追求高效生成与风格可控的视觉生成研究者

使用强化学习(RL)微调预训练生成模型已成为更贴近人类细微偏好的有效方法。本文研究将组相对策略优化(GRPO)应用于下一代视觉自回归(VAR)模型的微调。实验表明,该方法能有效对齐由美学预测器和CLIP嵌入生成的复杂奖励信号,显著提升图像质量,并实现对生成风格的精确控制。有趣的是,通过利用CLIP,该方法使VAR模型能够泛化至预训练阶段未包含的图像风格:在强化学习驱动的探索下,模型可生成与提示中提及的、但预训练中不存在的风格对齐的图像。综上所述,基于强化学习的微调对VAR模型既高效又有效,尤其得益于其快速推理速度,这对在线采样至关重要,而扩散模型在此方面面临显著挑战。

原文摘要 · Abstract (English)

Fine-tuning pre-trained generative models with Reinforcement Learning (RL) has emerged as an effective approach for aligning outputs more closely with nuanced human preferences. In this paper, we investigate the application of Group Relative Policy Optimization (GRPO) to fine-tune next-scale visual autoregressive (VAR) models. Our empirical results demonstrate that this approach enables alignment to intricate reward signals derived from aesthetic predictors and CLIP embeddings, significantly enhancing image quality and enabling precise control over the generation style. Interestingly, by leveraging CLIP, our method can help VAR models generalize beyond their initial ImageNet distribution: through RL-driven exploration, these models can generate images aligned with prompts referencing image styles that were absent during pre-training. In summary, we show that RL-based fine-tuning is both efficient and effective for VAR models, benefiting particularly from their fast inference speeds, which are advantageous for online sampling, an aspect that poses significant challenges for diffusion-based alternatives.

视觉生成强化学习风格控制VAR模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。