arXiv:2605.10730cs.CV2026-05被引 20

Qwen-Image-2.0统一生成与编辑,支持千字指令和多语言文本图像。

Qwen-Image-2.0 Technical Report

论文配图:Qwen-Image-2.0 Technical Report
图 1 · 摘自论文原文
  • 用Qwen3-VL编码条件,结合扩散模型实现联合建模。
  • 可生成1000词指令的图文内容,提升多语言文本清晰度。
  • 支持复杂场景下高保真图像生成,适合设计、创作人群使用。

我们提出Qwen-Image-2.0,一个统一高保真生成与精确编辑能力的通用图像生成基础模型。尽管已有进展,现有模型在超长文本渲染、多语言排版、高分辨率写实性、强指令遵循及高效部署方面仍面临挑战,尤其在文本丰富且构图复杂的场景中。Qwen-Image-2.0通过将Qwen3-VL作为条件编码器,并结合多模态扩散变换器实现条件-目标联合建模,辅以大规模数据清洗与定制化多阶段训练流程,既强化了多模态理解能力,又保持灵活生成与编辑性能。该模型支持高达1000个标记(tokens)的指令,用于生成幻灯片、海报、信息图、漫画等富文本内容,显著提升多语言文本保真度与排版质量。同时,在细节丰富度、真实纹理与光照一致性方面增强写实生成能力,对多样风格下的复杂提示更具鲁棒性。人工评估表明,相比先前版本,Qwen-Image-2.0在生成与编辑任务上均有显著提升,标志着迈向更通用、可靠、实用的图像生成基础模型的重要一步。

原文摘要 · Abstract (English)

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long text rendering, multilingual typography, high-resolution photorealism, robust instruction following, and efficient deployment, especially in text-rich and compositionally complex scenarios. Qwen-Image-2.0 addresses these challenges by coupling Qwen3-VL as the condition encoder with a Multimodal Diffusion Transformer for joint condition-target modeling, supported by large-scale data curation and a customized multi-stage training pipeline. This enables strong multimodal understanding while preserving flexible generation and editing capabilities. The model supports instructions of up to 1K tokens for generating text-rich content such as slides, posters, infographics, and comics, while significantly improving multilingual text fidelity and typography. It also enhances photorealistic generation with richer details, more realistic textures, and coherent lighting, and follows complex prompts more reliably across diverse styles. Extensive human evaluations show that Qwen-Image-2.0 substantially outperforms previous Qwen-Image models in both generation and editing, marking a step toward more general, reliable, and practical image generation foundation models.

图像生成多模态文本生成编辑能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。