arXiv:2509.20427cs.CV2025-09被引 236

一站式多模态图像生成,支持高效高分辨率创作与精准编辑

Seedream 4.0: Toward Next-generation Multimodal Image Generation

  • 统一文本到图像、编辑与多图合成的单框架设计
  • 2K图像生成仅需1.8秒,支持多图参考与复杂推理任务
  • 适合创意设计与专业内容生成,兼具速度与多模态能力

我们提出Seedream 4.0,一个高效且高性能的多模态图像生成系统,将文本到图像(T2I)生成、图像编辑与多图合成统一于单一框架。通过设计高效的扩散Transformer与强大VAE,显著减少图像令牌数量,实现高效训练,并可快速生成原生高分辨率图像(如1K-4K)。模型在数十亿跨领域、知识导向的图文对上预训练,结合数百个垂直场景的数据采集与优化策略,确保大规模稳定训练与强泛化能力。通过精心微调的视觉语言模型(VLM),联合进行多模态后训练,同时支持T2I与图像编辑任务。为加速推理,集成对抗蒸馏、分布匹配、量化及推测解码技术,可在不使用LLM/VLM作为外部模型时实现最高1.8秒生成2K图像。全面评估显示,Seedream 4.0在T2I与多模态图像编辑任务上均达当前最优性能,尤其在复杂任务中表现出色,包括精准编辑、上下文推理与多图参考生成。该系统将传统T2I工具扩展为更交互、多维度的创作平台,推动生成式AI在创意与专业应用中的边界。我们进一步推出扩展版本Seedream 4.5。Seedream 4.0与4.5可通过火山引擎体验:https://www.volcengine.com/experience/ark?launch=seedream。

原文摘要 · Abstract (English)

We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient diffusion transformer with a powerful VAE which also can reduce the number of image tokens considerably. This allows for efficient training of our model, and enables it to fast generate native high-resolution images (e.g., 1K-4K). Seedream 4.0 is pretrained on billions of text-image pairs spanning diverse taxonomies and knowledge-centric concepts. Comprehensive data collection across hundreds of vertical scenarios, coupled with optimized strategies, ensures stable and large-scale training, with strong generalization. By incorporating a carefully fine-tuned VLM model, we perform multi-modal post-training for training both T2I and image editing tasks jointly. For inference acceleration, we integrate adversarial distillation, distribution matching, and quantization, as well as speculative decoding. It achieves an inference time of up to 1.8 seconds for generating a 2K image (without a LLM/VLM as PE model). Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art results on both T2I and multimodal image editing. In particular, it demonstrates exceptional multimodal capabilities in complex tasks, including precise image editing and in-context reasoning, and also allows for multi-image reference, and can generate multiple output images. This extends traditional T2I systems into an more interactive and multidimensional creative tool, pushing the boundary of generative AI for both creativity and professional applications. We further scale our model and data as Seedream 4.5. Seedream 4.0 and Seedream 4.5 are accessible on Volcano Engine https://www.volcengine.com/experience/ark?launch=seedream.

图像生成多模态扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。