Qwen-Image提升复杂文本渲染与图像编辑一致性,支持中英文混合生成。
Qwen-Image Technical Report
- 构建全流程数据管道并采用渐进式训练,增强模型对复杂文本的生成能力。
- 在中英文文本渲染任务上均达顶尖水平,尤其在汉字等表意文字上表现突出。
- 双编码机制实现语义一致与视觉保真平衡,适合需要高精度图像编辑的场景。
我们提出Qwen-Image,是Qwen系列中的图像生成基础模型,在复杂文本渲染和精确图像编辑方面取得显著进展。为解决复杂文本渲染难题,设计了包含大规模数据收集、过滤、标注、合成与平衡的完整数据流水线,并采用从非文本到文本、由简至繁、逐步扩展至段落级描述的渐进式训练策略。该课程学习方法显著提升了模型的原生文本渲染能力。结果表明,Qwen-Image不仅在英语等拉丁字母语言上表现优异,还在更具挑战性的表意文字如中文上取得突破性进展。为提升图像编辑的一致性,引入改进的多任务训练范式,融合传统文本到图像(T2I)、文本图像到图像(TI2I)及图像到图像(I2I)重建任务,有效对齐了Qwen2.5-VL与MMDiT之间的潜在表示。同时,分别将原始图像输入Qwen2.5-VL和VAE编码器,获取语义与重构表征,实现双编码机制,使编辑模块在保持语义一致性与视觉保真度之间取得良好平衡。Qwen-Image在多个基准测试中达到当前最优性能,展现了强大的图像生成与编辑能力。
原文摘要 · Abstract (English)
We present Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. To address the challenges of complex text rendering, we design a comprehensive data pipeline that includes large-scale data collection, filtering, annotation, synthesis, and balancing. Moreover, we adopt a progressive training strategy that starts with non-text-to-text rendering, evolves from simple to complex textual inputs, and gradually scales up to paragraph-level descriptions. This curriculum learning approach substantially enhances the model's native text rendering capabilities. As a result, Qwen-Image not only performs exceptionally well in alphabetic languages such as English, but also achieves remarkable progress on more challenging logographic languages like Chinese. To enhance image editing consistency, we introduce an improved multi-task training paradigm that incorporates not only traditional text-to-image (T2I) and text-image-to-image (TI2I) tasks but also image-to-image (I2I) reconstruction, effectively aligning the latent representations between Qwen2.5-VL and MMDiT. Furthermore, we separately feed the original image into Qwen2.5-VL and the VAE encoder to obtain semantic and reconstructive representations, respectively. This dual-encoding mechanism enables the editing module to strike a balance between preserving semantic consistency and maintaining visual fidelity. Qwen-Image achieves state-of-the-art performance, demonstrating its strong capabilities in both image generation and editing across multiple benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。