8B参数模型媲美27B大模型,实现图像生成全流程统一
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

- 用统一变换器直接处理像素、文本和条件,告别分离的编码器和VAE
- 8B版本性能超越27B的Qwen-Image,200B+版本达新SOTA
- 适合追求高效统一架构的视觉生成研究与应用开发者
视觉生成模型长期受限于碎片化架构,依赖独立的文本编码器和外部VAE。本文提出HiDream-O1-Image,一种基于像素空间扩散Transformer的原生统一生成基础模型,开创从模块化架构到端到端上下文生成引擎的范式转变。通过将原始图像像素、文本标记和任务特定条件映射到单一共享标记空间,该模型在统一变换器(UiT)架构中实现多模态输入的结构统一。此原生编码范式消除了对独立VAE或分离预训练文本编码器的需求,使各类生成与编辑任务可视为一致的上下文推理过程。大量实验表明,HiDream-O1-Image在文本到图像生成、基于指令的编辑及主体驱动个性化等任务上表现优异。值得注意的是,仅80亿参数的版本即达到甚至超越参数量高达270亿的Qwen-Image模型性能。更重要的是,为验证该范式的巨大扩展性,模型成功扩展至超过2000亿参数。实验结果证明,2000亿以上参数版本的HiDream-O1-Image-Pro展现出前所未有的生成能力与卓越性能,树立了新的基准。最终,该工作揭示了原生统一架构的巨大潜力,并为下一代多模态AI指明了一条高度可扩展的发展路径。
原文摘要 · Abstract (English)
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDream-O1-Image, a natively unified generative foundation model via pixel-space Diffusion Transformer, that pioneers a paradigm shift from modular architectures to an end-to-end in-context visual generation engine. By mapping raw image pixels, text tokens, and task-specific conditions into a single shared token space, HiDream-O1-Image achieves a structural unification of multimodal inputs within an Unified Transformer (UiT) architecture. This native encoding paradigm eliminates the need for separate VAEs or disjoint pre-trained text encoders, allowing the model to treat diverse generation and editing tasks as a consistent in-context reasoning process. Extensive experiments show that HiDream-O1-Image excels across various generation tasks, including text-to-image generation, instruction-based editing, and subject-driven personalization. Notably, with only 8B parameters, HiDream-O1-Image (8B) achieves performance parity with or even surpasses established state-of-the-art models with significantly larger parameters (e.g., 27B Qwen-Image). Crucially, to validate the immense scalability of this paradigm, we successfully scale the architecture up to over 200B parameters. Experimental results demonstrate that this massive-scale version HiDream-O1-Image-Pro (200B+) unlocks unprecedented generative capabilities and superior performance, establishing new state-of-the-art benchmarks. Ultimately, HiDream-O1-Image highlights the immense potential of natively unified architectures and charts a highly scalable path toward next-generation multimodal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。