arXiv:2605.16842cs.AI2026-05

通过分阶段强化学习提升扩散多模态模型的图像生成质量

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

论文配图:Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
图 1 · 摘自论文原文
  • 分三阶段训练:先布局后结构再细化,匹配生成层级顺序
  • 在GenEval和DPG上显著提升图像质量与人类偏好评分
  • 针对关键结构令牌设计信用分配机制,更精准奖励贡献

扩散多模态大语言模型(dMLLMs)在图像生成中表现强大,但通过强化学习(RL)优化仍面临挑战。主要困难在于单张图像可通过多种未掩码序列生成,导致重要性比率计算往往不可行。此外,现有方法忽略dMLLM的层次化生成过程——早期令牌决定全局布局,后期令牌关注局部细节。统一奖励所有令牌的方法无法反映各令牌的真实贡献。为此,我们提出分层令牌GRPO(HT-GRPO),将生成层级直接融入策略优化。该方法采用‘草图-描画’训练方案,分三个阶段更新:全局、结构与精细化。同时引入提示条件估计器,从完全遮蔽状态计算重要性比率,并设计分层信用分配机制,优先激励关键结构令牌以确保奖励准确传播。在两个主流dMLLM骨干模型MMaDA和Lumina-DiMOO上的实验表明,HT-GRPO在GenEval和DPG基准上均取得显著提升。六项额外指标评估进一步证实其在图像质量、美学和人类偏好方面的显著改进。

原文摘要 · Abstract (English)

Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One primary difficulty is that a single image can be generated through many different unmasking sequences, which makes calculating importance ratios often intractable. Additionally, existing methods tend to ignore the hierarchical generation process of dMLLMs, where early tokens define the global layout and later tokens focus on local details. By assigning uniform rewards to all tokens, these current methods fail to reflect the actual contribution of each token to the final image. To address these issues, we propose Hierarchical Token GRPO (HT-GRPO), which integrates this hierarchy directly into the policy optimization process. Our approach features a Sketch-Then-Paint training scheme that organizes updates into three distinct stages: global, structure, and refinement. We also use a prompt-conditioned estimator to calculate importance ratios starting from a fully masked state. Furthermore, we introduce a Hierarchical Credit Assignment mechanism that prioritizes key structural tokens to ensure accurate reward propagation. Experiments using two popular dMLLM backbones, MMaDA and Lumina-DiMOO, demonstrate that HT-GRPO achieves substantial gains on the GenEval and DPG benchmarks. Evaluations across six additional metrics confirm significant improvements in image quality, aesthetics, and human preference.

扩散模型强化学习图像生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。