用强化学习提升离散自回归图像生成质量,实现图文统一建模。
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- 引入强化学习优化离散自回归图像生成过程
- 7B模型生成高审美质量图像并精准遵循复杂指令
- 适合需要图文一致生成的AI创作与交互系统
许多研究尝试将“下一个词预测”范式扩展至视觉内容,以建立图像生成与理解的统一方法。然而,通过离散令牌进行自回归建模的图像生成仍面临视觉保真度低、输出失真及难以遵循复杂指令等问题,可能源于自回归推理中的累积误差或离散化带来的信息损失。因此,近期研究更多转向联合训练扩散生成图像与自回归语言生成,偏离了统一建模。本文证明,强化学习可有效缓解伪影并显著提升离散自回归建模的生成质量,使图像与语言生成无缝融合。我们的框架X-Omni包含语义图像分词器、统一自回归模型(支持语言与图像)以及离线扩散解码器,使用7B语言模型在图像生成任务中达到当前最佳性能,生成图像具有高美学质量,同时具备强指令遵循能力与长文本渲染能力。
原文摘要 · Abstract (English)
Numerous efforts have been made to extend the ``next token prediction'' paradigm to visual contents, aiming to create a unified approach for both image generation and understanding. Nevertheless, attempts to generate images through autoregressive modeling with discrete tokens have been plagued by issues such as low visual fidelity, distorted outputs, and failure to adhere to complex instructions when rendering intricate details. These shortcomings are likely attributed to cumulative errors during autoregressive inference or information loss incurred during the discretization process. Probably due to this challenge, recent research has increasingly shifted toward jointly training image generation with diffusion objectives and language generation with autoregressive objectives, moving away from unified modeling approaches. In this work, we demonstrate that reinforcement learning can effectively mitigate artifacts and largely enhance the generation quality of a discrete autoregressive modeling method, thereby enabling seamless integration of image and language generation. Our framework comprises a semantic image tokenizer, a unified autoregressive model for both language and images, and an offline diffusion decoder for image generation, termed X-Omni. X-Omni achieves state-of-the-art performance in image generation tasks using a 7B language model, producing images with high aesthetic quality while exhibiting strong capabilities in following instructions and rendering long texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。