arXiv:2605.25347cs.CVcs.LG2026-05被引 2

ERNIE-Image是80亿参数的开源文生图模型,性能逼近商业顶尖水平。

ERNIE-Image Technical Report

论文配图:ERNIE-Image Technical Report
图 1 · 摘自论文原文
  • 采用分层数据构建策略,提升预训练数据质量与长尾概念覆盖
  • 在指令遵循、文字渲染和美学质量上超越现有开源模型,接近闭源顶尖水平
  • 提供轻量提示增强器与工业级美学评估基准,适合实际应用落地

我们介绍ERNIE-Image,一个基于80亿参数单流DiT架构的开源文生图模型。该模型通过更有效的大规模预训练数据挖掘和训练过程中的高质量监督,缩小开源模型与领先闭源系统间的差距。预训练阶段采用自下而上的数据构建流程,结合细粒度图像分类、丰富描述标注、美学评估与分层采样,降低数据噪声的同时保留长尾概念与真实世界细节知识,为复杂生成任务提供更强基础。后训练阶段采用自上而下的数据构建策略,针对高需求场景优化,多样化提示标注以匹配真实用户输入,并应用稳定化的DPO策略对齐人类审美偏好。我们进一步训练了支持8步推理的ERNIE-Image-Turbo,并提出MT-DMD方法缓解蒸馏中的能力退化问题。为提升实用性,引入轻量级提示增强器,将简洁用户意图扩展为结构化视觉描述。同时开发工业级美学模型ERNIE-Image-Aes及包含1000张人工标注的评估基准ERNIE-Image-Aes-1K。大量定性与定量实验表明,ERNIE-Image在开源模型中表现领先,其指令遵循、文本渲染与美学质量接近顶级商业模型。我们已公开训练模型与美学资源,以推动AIGC领域学术研究与技术进步。

原文摘要 · Abstract (English)

We introduce ERNIE-Image, an open-source text-to-image generation model built upon an 8B single-stream DiT architecture. ERNIE-Image aims to bridge the gap between current open-source models and leading closed-source systems through more effective mining of large-scale pre-training data and improved supervision quality throughout training. During pre-training, we adopt a bottom-up data construction pipeline that combines fine-grained image categorization, rich caption annotation, aesthetic assessment, and hierarchical sampling. This strategy reduces data noise while preserving long-tail concepts and detailed real-world knowledge, providing a stronger foundation for complex generation tasks. In the post-training stage, we use a top-down data construction pipeline for high-demand scenarios, diversify prompt annotations to better match real user inputs, and apply a stabilized DPO strategy to align the model with human aesthetic preferences. We further train ERNIE-Image-Turbo for efficient 8-NFE generation and propose MT-DMD to mitigate capability drift during distillation. To make the model easier to use in practical scenarios, we equip it with a lightweight Prompt Enhancer that expands concise user intents into structured visual descriptions. In addition, we develop ERNIE-Image-Aes, an industrial-grade aesthetic model, together with ERNIE-Image-Aes-1K, a human-annotated benchmark for realistic aesthetic evaluation. Extensive qualitative and quantitative experiments show that ERNIE-Image achieves leading performance among open-source models and approaches top-tier commercial models in instruction following, text rendering, and aesthetic quality. We release the trained models and aesthetic resources to facilitate further academic research and technical progress in the AIGC community.

文生图扩散模型开源模型美学评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。