arXiv:2504.21356cs.CVcs.AI2025-04被引 10

统一图像理解生成编辑,用预填充自回归提升质量与可控性

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

  • 共享嵌入空间连接自回归与扩散模型,融合跨模态优势
  • 引入预填充自回归策略,缓解生成过程中的误差累积问题
  • 在2630万样本数据上训练,三任务均达当前最佳性能

统一多模态生成模型旨在整合图像理解与生成能力,尤其在交错文本-图像数据上具有显著优势。然而现有统一模型存在图像合成质量低、自回归误差累积严重、图像编辑能力弱等局限。本文提出Nexus-Gen,一种在共享图像嵌入空间中统一图像理解、生成与编辑任务的新架构。该共享空间作为自回归模型与扩散模型的桥梁,无缝融合二者在跨模态建模中的互补优势。为缓解自回归嵌入预测中的严重误差累积,我们提出新颖的预填充自回归策略,通过可学习嵌入预先填充输入序列,使训练与推理动态对齐。在包含2630万样本的大规模自建数据集上进行多阶段多任务训练后,Nexus-Gen在涵盖图像理解、生成与编辑任务的评估基准上均取得当前最优表现。所有模型、数据集及源代码已开源至https://github.com/modelscope/Nexus-Gen,以促进领域进一步发展。

原文摘要 · Abstract (English)

Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly interleaved text-image data. However, existing unified models exhibit limitations in image synthesis quality, autoregressive error accumulation, and image editing capability. In this work, we propose Nexus-Gen, a novel architecture that unifies image understanding, generation, and editing tasks in a shared image embedding space. This shared space serves as a bridge for the autoregressive and diffusion models, which seamlessly integrates their complementary strengths in cross-modal modeling. To mitigate the severe error accumulation during autoregressive embedding prediction, we propose a novel prefilled autoregression strategy that aligns training-inference dynamics by prefilling input sequences with learnable embeddings. After multi-stage and multi-task training on our constructed large-scale dataset with 26.3 million samples, Nexus-Gen achieves state-of-the-art performance on the evaluation benchmarks spanning image understanding, generation and editing tasks. All models, datasets, and source codes are released in https://github.com/modelscope/Nexus-Gen to facilitate further advancements across the field.

图像生成自回归多任务扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。