让AI生成图具备专业级可控性,支持复杂文字与多主体保真
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

- 融合大语言模型与扩散模型,实现意图到图像的精准转化
- 支持超长文本渲染、多主体身份保留、4K高清高效生成
- 适合电商、设计、教育等需要高精度视觉创作的场景
我们提出Wan-Image,一个统一的视觉生成系统,旨在将图像生成模型从随意合成工具转变为专业生产力工具。当前扩散模型虽在美学生成上表现优异,但在需绝对可控性、复杂排版和严格身份保持的设计流程中常遇瓶颈。Wan-Image通过原生统一的多模态架构,结合大语言模型的认知能力与扩散变换器的高保真像素合成能力,无缝将细微用户意图转化为精确视觉输出。其基础依赖大规模多模态数据扩展、细粒度标注引擎及精选强化学习数据,超越基础指令跟随,实现专家级能力:包括超长复杂文本渲染、高度多样肖像生成、调色板引导生成、多主体身份保留、连贯序列视觉生成、精确多模态交互编辑、原生透明通道生成及高效4K合成。在多项人类评估中,Wan-Image整体性能优于Seedream 5.0 Lite和GPT Image 1.5,挑战任务下与Nano Banana Pro相当。最终,它推动电商、娱乐、教育和个人生产力领域的视觉内容创作革新,重新定义专业视觉合成边界。
原文摘要 · Abstract (English)
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。