让视觉理解和生成协同进化,实现智能图像生成
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- 通过两阶段训练,使模型在生成中具备真实思维链
- 强化学习提升生成质量,实现图文任务统一
- 适合研究多模态生成与推理的学者
当前多模态大语言模型(MLLMs)虽致力于统一视觉理解与生成,但两者仍相互独立,如同嵌入同一模型的两个分离功能。因此,视觉理解无法促进生成,且大语言模型的推理机制也未充分融入图像生成过程。本文提出一种协同共进化机制,推动视觉理解与生成的深度融合,将图像生成变为迭代内省过程。我们采用两阶段训练:监督微调赋予模型生成真实思维链(CoT)的基础能力;强化学习通过探索-利用权衡激发其全部潜力。最终,模型实现生成中的‘顿悟时刻’,使MLLM从文本到图像任务跃升为统一的图像生成系统。大量实验表明,该模型不仅在文本到图像生成和图像编辑上表现优异,还具备更强的视觉理解能力,可作为高效的图像语义评估器。项目页面:https://janus-pro-r1.github.io。
原文摘要 · Abstract (English)
Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same model. Consequently, visual comprehension does not enhance visual generation, and the reasoning mechanisms of LLMs have not been fully integrated to revolutionize image generation. In this paper, we propose to enable the collaborative co-evolution of visual comprehension and generation, advancing image generation into an iterative introspective process. We introduce a two-stage training approach: supervised fine-tuning teaches the MLLM with the foundational ability to generate genuine CoT for visual generation, while reinforcement learning activates its full potential via an exploration-exploitation trade-off. Ultimately, we unlock the Aha moment in visual generation, advancing MLLMs from text-to-image tasks to unified image generation. Extensive experiments demonstrate that our model not only excels in text-to-image generation and image editing, but also functions as a superior image semantic evaluator with enhanced visual comprehension capabilities. Project Page: https://janus-pro-r1.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。