arXiv:2510.15857cs.CV2025-10被引 49

BLIP3o-NEXT统一图文生成与编辑,实现更真实连贯的图像生成。

BLIP3o-NEXT: Next Frontier of Native Image Generation

  • 采用自回归+扩散模型架构,结合推理与细节渲染优势。
  • 在多个图文生成与编辑任务中超越现有模型表现。
  • 适合关注图像生成前沿与开源模型的开发者与研究者。

我们提出BLIP3o-NEXT,一个完全开源的BLIP3系列基础模型,推动原生图像生成的新前沿。该模型将文本到图像生成与图像编辑统一于单一架构中,展现出强大的生成与编辑能力。在构建顶尖原生图像生成模型的过程中,我们总结出四大关键洞察:(1) 多数架构选择性能相近;只要具备高效扩展性与快速推理能力,即视为有效;(2) 强化学习的成功应用可进一步拓展原生图像生成边界;(3) 图像编辑仍是挑战,但通过后训练与数据工程可显著提升指令遵循能力与生成图像与参考图的一致性;(4) 数据质量与规模仍是决定模型性能上限的关键因素。基于这些洞察,BLIP3o-NEXT采用自回归+扩散架构:自回归模型首先根据多模态输入生成离散图像标记,其隐藏状态作为条件信号驱动扩散模型生成高保真图像。该架构融合自回归模型的推理能力与指令跟随优势,以及扩散模型的精细渲染能力,实现了更高层次的一致性与真实感。在多种文本到图像与图像编辑基准上的广泛评估表明,BLIP3o-NEXT在性能上优于现有模型。

原文摘要 · Abstract (English)

We present BLIP3o-NEXT, a fully open-source foundation model in the BLIP3 series that advances the next frontier of native image generation. BLIP3o-NEXT unifies text-to-image generation and image editing within a single architecture, demonstrating strong image generation and image editing capabilities. In developing the state-of-the-art native image generation model, we identify four key insights: (1) Most architectural choices yield comparable performance; an architecture can be deemed effective provided it scales efficiently and supports fast inference; (2) The successful application of reinforcement learning can further push the frontier of native image generation; (3) Image editing still remains a challenging task, yet instruction following and the consistency between generated and reference images can be significantly enhanced through post-training and data engine; (4) Data quality and scale continue to be decisive factors that determine the upper bound of model performance. Building upon these insights, BLIP3o-NEXT leverages an Autoregressive + Diffusion architecture in which an autoregressive model first generates discrete image tokens conditioned on multimodal inputs, whose hidden states are then used as conditioning signals for a diffusion model to generate high-fidelity images. This architecture integrates the reasoning strength and instruction following of autoregressive models with the fine-detail rendering ability of diffusion models, achieving a new level of coherence and realism. Extensive evaluations of various text-to-image and image-editing benchmarks show that BLIP3o-NEXT achieves superior performance over existing models.

图像生成扩散模型自回归开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。