arXiv:2510.02880cs.AI2025-10被引 8

提出MaskGRPO,让离散扩散模型在多模态任务中实现高效强化学习。

Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

  • 设计重要性采样机制,解决非自回归结构下的梯度更新难题
  • 在数学推理、编程和视觉生成任务上性能显著提升
  • 适合需要高质量多模态生成与稳定优化的研究者

用奖励优化离散扩散模型(DDM)仍具挑战:其非自回归结构导致重要性采样不可行且轨迹计算复杂,使群组相对策略优化(GRPO)等强化学习方法难以应用。本文提出MaskGRPO,首个可扩展的多模态强化学习方法,具备有效重要性采样和模态特异性适应能力。首先厘清DDM的理论基础,构建捕捉关键标记波动的梯度估计器;随后针对视觉序列定制滚动策略,生成多样化补全并提供可靠优化梯度。在数学推理、代码生成与视觉生成基准测试中,MaskGRPO带来更稳定高效的更新,提升推理性能与生成质量。本研究确立了MaskGRPO为系统性策略优化方法,也是首个实用的离散化视觉扩散强化学习方案。

原文摘要 · Abstract (English)

Optimizing discrete diffusion model (DDM) with rewards remains a challenge: the non-autoregressive paradigm makes importance sampling intractable and rollout complex, puzzling reinforcement learning methods such as Group Relative Policy Optimization (GRPO). In this study, we introduce MaskGRPO, the first viable approach to enable scalable multimodal reinforcement learning in discrete diffusion with effective importance sampling and modality-specific adaptations. To this end, we first clarify the theoretical foundation for DDMs, which facilitates building an importance estimator that captures valuable token fluctuation for gradient updates. We then delicately tailored the rollout method for visual sequences, which yields diverse completions and reliable optimization gradients. Upon math reasoning, coding, and visual generation benchmarks, MaskGRPO brings more stable and efficient updates, leading to stronger reasoning performance and better generation quality. This study establishes MaskGRPO as a systematic policy optimization approach and the first practical way for discretized visual diffusion.

强化学习扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。