用二值隐码统一生成与识别任务,提升图像生成与表征能力。
BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
- 通过二值编码器和掩码建模,实现生成与判别一体化。
- 在FID-50k上表现优于现有模型,线性探测准确率达92.3%。
- 零样本泛化支持图像修复、编辑等应用,无需结构改动。
我们提出BiGR,一种基于紧凑二值隐码的新型条件图像生成模型,旨在同时增强生成质量与视觉表征能力。BiGR是首个在统一框架内融合生成与判别任务的条件生成模型,包含二值分词器、掩码建模机制和二值译码器,用于预测二值代码。此外,我们引入熵有序采样方法以实现高效图像生成。大量实验验证了BiGR在生成质量(以FID-50k衡量)和表征能力(以线性探测准确率评估)上的卓越性能。同时,其展现出跨多种视觉任务的零样本泛化能力,支持图像修复、外推、编辑、插值与增强等应用,且无需结构修改。研究结果表明,BiGR能有效统一生成与判别任务,为该领域未来发展提供新路径。我们还使BiGR具备文本到图像生成能力,进一步拓展其应用潜力。
原文摘要 · Abstract (English)
We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same framework. BiGR features a binary tokenizer, a masked modeling mechanism, and a binary transcoder for binary code prediction. Additionally, we introduce a novel entropy-ordered sampling method to enable efficient image generation. Extensive experiments validate BiGR's superior performance in generation quality, as measured by FID-50k, and representation capabilities, as evidenced by linear-probe accuracy. Moreover, BiGR showcases zero-shot generalization across various vision tasks, enabling applications such as image inpainting, outpainting, editing, interpolation, and enrichment, without the need for structural modifications. Our findings suggest that BiGR unifies generative and discriminative tasks effectively, paving the way for further advancements in the field. We further enable BiGR to perform text-to-image generation, showcasing its potential for broader applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。