arXiv:2511.14760cs.CV2025-11被引 7

用统一奖励机制提升图像生成与编辑能力,性能接近闭源模型。

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

  • 通过共享奖励模型的强化学习策略,统一优化生成与编辑任务。
  • 在GenEval和ImgEdit评测中得分分别达0.89和4.31,超越现有开源模型。
  • 轻量级指令对齐阶段显著提升编辑理解力,助力强化学习训练。

我们提出UniGen-1.5,一种用于先进图像理解、生成与编辑的统一多模态大语言模型。在UniGen基础上,全面优化模型架构与训练流程,增强图像理解与生成能力,并解锁强大的图像编辑功能。特别地,我们提出一种统一的强化学习(RL)策略,通过共享奖励模型联合提升图像生成与编辑性能。为进一步提升编辑表现,我们设计了一个轻量级编辑指令对齐阶段,显著增强对编辑指令的理解能力,这对强化学习训练至关重要。实验结果表明,UniGen-1.5在图像理解与生成方面表现具有竞争力。具体而言,在GenEval和ImgEdit评测中,其综合得分分别为0.89和4.31,超越BAGEL等当前最优开源模型,达到GPT-Image-1等专有模型的水平。

原文摘要 · Abstract (English)

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen the image understanding and generation capabilities while unlocking strong image editing ability. Especially, we propose a unified Reinforcement Learning (RL) strategy that improves both image generation and image editing jointly via shared reward models. To further enhance image editing performance, we propose a light Edit Instruction Alignment stage that significantly improves the editing instruction comprehension that is essential for the success of the RL training. Experimental results show that UniGen-1.5 demonstrates competitive understanding and generation performance. Specifically, UniGen-1.5 achieves 0.89 and 4.31 overall scores on GenEval and ImgEdit that surpass the state-of-the-art models such as BAGEL and reaching performance comparable to proprietary models such as GPT-Image-1.

图像生成强化学习多模态编辑能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。