arXiv:2605.12495cs.CVcs.AI2026-05被引 2

让多模态模型自己反思并改进生成结果,提升准确率。

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

论文配图:AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
图 1 · 摘自论文原文
  • 用分解式可验证奖励机制,让模型分步评估生成质量。
  • 在多个基准上表现优于现有方法,尤其在理解隐含意图时更准。
  • 适合需要高精度生成与自动纠错的AI应用开发者。

本文提出AlphaGRPO,一种将组相对策略优化(GRPO)应用于AR-Diffusion统一多模态模型(UMMs)的新框架,无需冷启动阶段即可增强多模态生成能力。该方法使模型具备高级推理能力:主动推断用户隐含意图的文本到图像生成,以及自主诊断并修正生成内容偏差的自反思优化。为应对真实场景中稳定监督的挑战,我们引入分解式可验证奖励(DVReward)——利用大语言模型将复杂请求分解为原子级、可验证的语义与质量问题,再由通用多模态大模型评估,提供可靠且可解释的反馈。大量实验表明,AlphaGRPO在GenEval、TIIF-Bench、DPG-Bench和WISE等多模态生成基准上均取得显著提升,并在无编辑训练的情况下,在GEdit数据集上实现编辑任务性能跃升。结果验证了自反思强化学习能有效利用模型内在理解能力,引导高保真生成。

原文摘要 · Abstract (English)

In this paper, we propose AlphaGRPO, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models (UMMs) to enhance multimodal generation capabilities without an additional cold-start stage. Our approach unlocks the model's intrinsic potential to perform advanced reasoning tasks: Reasoning Text-to-Image Generation, where the model actively infers implicit user intents, and Self-Reflective Refinement, where it autonomously diagnoses and corrects misalignments in generated outputs. To address the challenge of providing stable supervision for real-world multimodal generation, we introduce the Decompositional Verifiable Reward (DVReward). Unlike holistic scalar rewards, DVReward utilizes an LLM to decompose complex user requests into atomic, verifiable semantic and quality questions, which are then evaluated by a general MLLM to provide reliable and interpretable feedback. Extensive experiments demonstrate that AlphaGRPO yields robust improvements across multimodal generation benchmarks, including GenEval, TIIF-Bench, DPG-Bench and WISE, while also achieving significant gains in editing tasks on GEdit without training on editing tasks. These results validate that our self-reflective reinforcement approach effectively leverages inherent understanding to guide high-fidelity generation. Project page: https://huangrh99.github.io/AlphaGRPO/

多模态生成自反思强化学习提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。