统一视觉语言模型,用循环一致性提升图像生成与理解能力。
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models
- 通过图像→布局→图像的循环训练,实现生成与理解一体化。
- 在多个基准上表现优于独立模块模型,自监督生成效果显著提升。
- 适合研究统一多模态模型、生成式AI的开发者与学者。
我们提出CyCLeGen,一种统一的视觉-语言基础模型,可在单一自回归框架内完成图像理解与生成。不同于依赖独立感知与合成模块的现有模型,CyCLeGen采用全集成架构,通过图像→布局→图像和布局→图像→布局的循环一致性学习进行训练。该统一框架带来两大优势:自我反思能力,使模型能推理自身生成结果;数据效率高,可在强化学习目标引导下,基于循环一致性实现自生成监督下的自我优化。大量实验表明,CyCLeGen在多样化的图像理解与生成基准上均取得显著提升,凸显了统一视觉-语言基础模型的巨大潜力。
原文摘要 · Abstract (English)
We present CyCLeGen, a unified vision-language foundation model capable of both image understanding and image generation within a single autoregressive framework. Unlike existing vision models that depend on separate modules for perception and synthesis, CyCLeGen adopts a fully integrated architecture that enforces cycle-consistent learning through image->layout->image and layout->image->layout generation loops. This unified formulation introduces two key advantages: introspection, enabling the model to reason about its own generations, and data efficiency, allowing self-improvement via synthetic supervision under a reinforcement learning objective guided by cycle consistency. Extensive experiments show that CyCLeGen achieves significant gains across diverse image understanding and generation benchmarks, highlighting the potential of unified vision-language foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。