arXiv:2504.11346cs.CV2025-04被引 145

Seedream 3.0 提升中英双语图像生成质量与速度,支持高分辨率文字渲染。

Seedream 3.0 Technical Report

  • 全链路优化数据构建到部署,提升多语言图文对齐能力。
  • 支持最高2K分辨率,中文复杂字符生成更精准,视觉质量显著提升。
  • 采用新型加速策略,生成速度提升4至8倍,适合专业设计场景。

我们提出 Seedream 3.0,一个高性能中英文双语图像生成基础模型。针对 Seedream 2.0 存在的提示对齐困难、细粒度文字生成不足、视觉美感与保真度欠佳、图像分辨率受限等问题,从数据构建到模型部署全链路进行改进。数据层面,采用缺陷感知训练范式与双轴协同采样框架,数据量翻倍;预训练阶段引入混合分辨率训练、跨模态 RoPE、表示对齐损失及分辨率感知时间步采样。后训练阶段使用多样化审美标注进行 SFT,结合基于 VLM 的带缩放奖励模型,实现与人类偏好高度对齐。此外,首创一致性噪声期望与重要性感知时间步采样加速范式,实现4至8倍速度提升,同时保持图像质量。相比 Seedream 2.0,Seedream 3.0 显著增强整体能力,尤其在复杂中文字符的文字渲染方面表现突出,支持原生2K高分辨率输出,生成图像视觉质量优异。

原文摘要 · Abstract (English)

We present Seedream 3.0, a high-performance Chinese-English bilingual image generation foundation model. We develop several technical improvements to address existing challenges in Seedream 2.0, including alignment with complicated prompts, fine-grained typography generation, suboptimal visual aesthetics and fidelity, and limited image resolutions. Specifically, the advancements of Seedream 3.0 stem from improvements across the entire pipeline, from data construction to model deployment. At the data stratum, we double the dataset using a defect-aware training paradigm and a dual-axis collaborative data-sampling framework. Furthermore, we adopt several effective techniques such as mixed-resolution training, cross-modality RoPE, representation alignment loss, and resolution-aware timestep sampling in the pre-training phase. During the post-training stage, we utilize diversified aesthetic captions in SFT, and a VLM-based reward model with scaling, thereby achieving outputs that well align with human preferences. Furthermore, Seedream 3.0 pioneers a novel acceleration paradigm. By employing consistent noise expectation and importance-aware timestep sampling, we achieve a 4 to 8 times speedup while maintaining image quality. Seedream 3.0 demonstrates significant improvements over Seedream 2.0: it enhances overall capabilities, in particular for text-rendering in complicated Chinese characters which is important to professional typography generation. In addition, it provides native high-resolution output (up to 2K), allowing it to generate images with high visual quality.

图像生成双语模型高分辨率文本渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。