arXiv:2503.07703cs.CV2025-03被引 66

Seedream 2.0 支持中英文双语生成,更懂中国文化细节。

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

  • 自研双语大模型作文本编码器,原生融合中英知识
  • 多阶段优化使图文一致性和美学表现达顶尖水平
  • 适合需要精准中文化表达的图像生成与编辑场景

扩散模型的快速发展推动了图像生成技术的飞跃。然而,主流模型如Flux、SD3.5和Midjourney仍存在模型偏见、文本渲染能力弱及对中国文化理解不足等问题。为此,我们提出Seedream 2.0,一个原生支持中英文双语的图像生成基础模型,在中英文提示处理、双语图像生成与文本渲染方面表现出色。通过构建强大的数据系统与平衡准确与丰富的描述标注体系,并集成自研双语大语言模型作为文本编码器,实现从海量数据中直接学习原生知识。模型采用Glyph-Aligned ByT5实现灵活字符级文本渲染,结合可扩展的ROPE适配未训练分辨率。经过多阶段后训练优化(包括SFT与RLHF迭代),在提示遵循性、美学质量、文本渲染与结构正确性上均达到领先水平。实验表明,其输出在人类偏好评估中获得优异ELO分数。此外,它可轻松适配为指令式图像编辑模型SeedEdit,具备强指令遵循与图像一致性平衡能力。

原文摘要 · Abstract (English)

Rapid advancement of diffusion models has catalyzed remarkable progress in the field of image generation. However, prevalent models such as Flux, SD3.5 and Midjourney, still grapple with issues like model bias, limited text rendering capabilities, and insufficient understanding of Chinese cultural nuances. To address these limitations, we present Seedream 2.0, a native Chinese-English bilingual image generation foundation model that excels across diverse dimensions, which adeptly manages text prompt in both Chinese and English, supporting bilingual image generation and text rendering. We develop a powerful data system that facilitates knowledge integration, and a caption system that balances the accuracy and richness for image description. Particularly, Seedream is integrated with a self-developed bilingual large language model as a text encoder, allowing it to learn native knowledge directly from massive data. This enable it to generate high-fidelity images with accurate cultural nuances and aesthetic expressions described in either Chinese or English. Beside, Glyph-Aligned ByT5 is applied for flexible character-level text rendering, while a Scaled ROPE generalizes well to untrained resolutions. Multi-phase post-training optimizations, including SFT and RLHF iterations, further improve the overall capability. Through extensive experimentation, we demonstrate that Seedream 2.0 achieves state-of-the-art performance across multiple aspects, including prompt-following, aesthetics, text rendering, and structural correctness. Furthermore, Seedream 2.0 has been optimized through multiple RLHF iterations to closely align its output with human preferences, as revealed by its outstanding ELO score. In addition, it can be readily adapted to an instruction-based image editing model, such as SeedEdit, with strong editing capability that balances instruction-following and image consistency.

图像生成双语模型文化理解文本渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。