用视觉语言模型生成图像,无需复杂训练就能高保真合成。
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

- 用可学习查询令牌从预训练VLM提取语义嵌入作为扩散模型条件。
- 在多个图文生成与编辑任务中优于当前最佳基线模型。
- 适合需要高效多模态生成且不想从头训练的研究者。
我们提出MMCORE,一个统一的多模态图像生成与编辑框架。MMCORE利用预训练的视觉-语言模型(VLM)通过可学习查询令牌预测语义视觉嵌入,并将其作为扩散模型的条件信号。该设计有效将VLM丰富的理解与推理能力迁移至视觉生成过程。通过避免自回归与扩散模型间的深层融合或从零训练,MMCORE显著降低计算开销,同时保持高质量合成效果。MMCORE无缝集成文本到图像生成与交错式图像生成,在空间推理和视觉定位等复杂场景下展现出强大的多模态理解能力。全面评估表明,MMCORE在广泛的文字转图像及单/多图像编辑基准测试中持续优于现有最先进方法。
原文摘要 · Abstract (English)
We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, which subsequently serve as conditioning signals for a diffusion model. This streamlined design effectively transfers the rich understanding and reasoning capabilities of VLMs into the visual generation process. By obviating the need for deep fusion between autoregressive and diffusion models or training from scratch, MMCORE significantly reduces computational overhead while maintaining high-fidelity synthesis. MMCORE seamlessly integrates text-to-image synthesis with interleaved image generation, demonstrating robust multimodal comprehension in complex scenarios such as spatial reasoning and visual grounding. Comprehensive evaluations indicate that MMCORE consistently outperforms state-of-the-art baselines across a broad spectrum of text-to-image and single/multi-image editing benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。