arXiv:2603.06043cs.CV2026-03中稿 · CVPR被引 3

用理解能力提升生成质量,让模型自己教自己画得更好

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

  • 用理解分支评估生成结果,反向优化生成过程
  • 在多个数据集上显著提升图像生成质量,最高提升18.7%
  • 无需外部标注,适合想改进生成能力的开发者

统一多模态模型(UMMs)在视觉理解与生成融合方面取得显著进展,展现出处理复杂文本到图像(T2I)任务的强大潜力。然而,其理论优势与其实际表现之间存在持续的能力差距:尽管理解能力出色,但生成能力相对薄弱。这主要源于理解与生成过程的内在脱节。为解决该问题,我们利用模型内部的理解能力来增强生成质量。提出一种基于标记级的内在图文对齐奖励机制GvU,使模型能同时作为教师和学生:通过理解分支评估自身生成结果,并据此引导生成优化。在此基础上,构建自监督强化学习框架,使UMMs能通过理解驱动的内在奖励信号迭代改进生成质量,无需依赖外部监督。实验表明,该方法显著提升了生成质量,反过来也增强了细粒度视觉理解能力,缩小了统一多模态模型在理解与生成间的性能差距。

原文摘要 · Abstract (English)

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a persistent capability gap exists: UMMs typically exhibit superior visual understanding but comparatively weaker generative capabilities. This discrepancy arises largely from the intrinsic decoupling between the understanding and generation processes. While a UMM can accurately interpret fine-grained visual details, it often struggles to produce semantically coherent images from complex textual prompts. To address this challenge, we explore UMMs' internal understanding capability to enhance generation quality. We propose a token-level intrinsic text-image alignment reward mechanism, GvU, enabling the UMM to act simultaneously as teacher and student: it evaluates its own outputs using the understanding branch to guide the generations accordingly. Building upon this, we design a self-supervised reinforcement learning framework, allowing UMMs to iteratively improve their generation quality through understanding-based intrinsic reward signals--without reliance on external supervision. Experimental results show that our method substantially boosts UMMs' generation, which in turn strengthens their fine-grained visual understanding, narrowing the capability gap between UMMs' visual understanding and generation.

多模态生成自监督学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。