高压缩图像生成新模型,重建和扩散能力双突破。
Qwen-Image-VAE-2.0 Technical Report

- 采用全局跳跃连接与扩展隐空间,提升高压缩下的重建质量。
- 在真实文档场景中实现顶尖重建效果,且支持快速扩散建模。
- 适合需要高效图像压缩与高质量生成的视觉应用开发者。
我们提出 Qwen-Image-VAE-2.0,一套高压缩率的变分自编码器,显著提升重建保真度与扩散建模能力。为解决高压缩下的重建瓶颈,采用改进架构,包含全局跳跃连接(GSC)与扩展的隐层通道数;通过百亿级图像训练并引入合成渲染引擎,增强文本密集场景性能。针对高维隐空间收敛难题,设计增强语义对齐策略,使隐空间更适配扩散模型。为优化计算效率,使用非对称无注意力编码器-解码器结构以降低编码开销。我们在公开重建基准上进行全面评估,并提出 OmniDoc-TokenBench 基准,包含多样真实文档及基于 OCR 的专用评估指标,用于评测文本丰富场景表现。Qwen-Image-VAE-2.0 在高压缩比下达成当前最优重建性能,尤其在通用领域与文本密集场景均表现卓越。下游 DiT 实验显示其具备优异扩散能力,收敛速度远超现有高压缩基线。这确立了 Qwen-Image-VAE-2.0 作为兼具高压缩、高保真与强扩散性的领先模型地位。
原文摘要 · Abstract (English)
We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To address the reconstruction bottlenecks of high compression, we adopt an improved architecture featuring Global Skip Connections (GSC) and expanded latent channels. Moreover, we scale training to billions of images and incorporate a synthetic rendering engine to improve performance in text-rich scenarios. To tackle the convergence challenges of high-dimensional latent space, we implement an enhanced semantic alignment strategy to make the latent space highly amenable to diffusion modeling. To optimize computational efficiency, we leverage an asymmetric and attention-free encoder-decoder backbone to minimize encoding overhead. We present a comprehensive evaluation of Qwen-Image-VAE-2.0 on public reconstruction benchmarks. To evaluate performance in text-rich scenarios, we propose OmniDoc-TokenBench, a new benchmark comprising a diverse collection of real-world documents coupled with specialized OCR-based evaluation metrics. Qwen-Image-VAE-2.0 achieves state-of-the-art reconstruction performance, demonstrating exceptional capabilities in both general domains and text-rich scenarios at high compression ratio. Furthermore, downstream DiT experiments reveal our models possess superior diffusability, significantly accelerating convergence compared to existing high-compression baselines. These establish Qwen-Image-VAE-2.0 as a leading model with high compression, superior reconstruction, and exceptional diffusability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。