针对自回归图像生成,提出聚焦关键帧的优化方法,提升生成质量。
Group Critical-token Policy Optimization for Autoregressive Image Generation
- 从因果依赖、熵梯度结构、奖励感知多样性三方面识别关键图像标记
- 仅用30%关键标记即超越全标记优化方法,性能更优
- 适合追求高效高质图像生成的研究者与开发者
近期研究将可验证奖励强化学习(RLVR)拓展至自回归(AR)视觉生成并取得显著进展。然而,现有方法对所有图像标记采用统一优化策略,未考虑不同标记在RLVR训练中的贡献差异。核心挑战在于如何识别生成过程中更具影响力的图像标记,并对其实施有效的逐标记优化。为此,我们提出组关键标记策略优化(GCPO),通过三个维度识别关键标记:(1) 因果依赖性——早期标记因单向依赖决定后续标记与最终图像效果;(2) 熵诱导的空间结构——高熵梯度标记对应图像结构及区域间连接;(3) 基于RLVR的标记多样性——一组采样图像中视觉相似度低的标记有助于提升标记级多样性。针对这些关键标记,引入动态逐标记优势权重,基于策略模型与参考模型之间的置信度差异促进探索。仅使用30%图像标记的GCPO,在多个文本到图像基准测试中,包括AR模型与统一多模态模型,均优于使用全部标记的GRPO,展现出卓越的有效性。
原文摘要 · Abstract (English)
Recent studies have extended Reinforcement Learning with Verifiable Rewards (RLVR) to autoregressive (AR) visual generation and achieved promising progress. However, existing methods typically apply uniform optimization across all image tokens, while the varying contributions of different image tokens for RLVR's training remain unexplored. In fact, the key obstacle lies in how to identify more critical image tokens during AR generation and implement effective token-wise optimization for them. To tackle this challenge, we propose $\textbf{G}$roup $\textbf{C}$ritical-token $\textbf{P}$olicy $\textbf{O}$ptimization ($\textbf{GCPO}$), which facilitates effective policy optimization on critical tokens. We identify the critical tokens in RLVR-based AR generation from three perspectives, specifically: $\textbf{(1)}$ Causal dependency: early tokens fundamentally determine the later tokens and final image effect due to unidirectional dependency; $\textbf{(2)}$ Entropy-induced spatial structure: tokens with high entropy gradients correspond to image structure and bridges distinct visual regions; $\textbf{(3)}$ RLVR-focused token diversity: tokens with low visual similarity across a group of sampled images contribute to richer token-level diversity. For these identified critical tokens, we further introduce a dynamic token-wise advantage weight to encourage exploration, based on confidence divergence between the policy model and reference model. By leveraging 30\% of the image tokens, GCPO achieves better performance than GRPO with full tokens. Extensive experiments on multiple text-to-image benchmarks for both AR models and unified multimodal models demonstrate the effectiveness of GCPO for AR visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。