arXiv:2504.01934cs.CV2025-04被引 61

ILLUME+通过双视觉分词与扩散解码,实现理解、生成、编辑三者统一

ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

  • 采用双视觉分词器保留纹理与语义,支持粗到细的多模态表征
  • 用扩散模型解码提升图像生成质量,实现高效超分辨率
  • 3B参数模型在理解、生成、编辑任务中均表现优异,适合多场景应用

我们提出 ILLUME+,通过双视觉分词与扩散解码器,同时提升深层语义理解与高保真图像生成能力。现有统一模型难以兼顾理解、生成与编辑三大核心能力:如 Chameleon 和 EMU3 使用 VQGAN 进行图像离散化,因缺乏深层语义交互,在视觉理解任务上落后于 LLaVA 等专用模型;而 LaViT 与 ILLUME 虽采用语义编码器分词,却在图像编辑中纹理保留差;Janus 系列则分离输入输出图像表示,限制了图文交错任务的无缝处理。相比之下,ILLUME+ 引入统一的双视觉分词器 DualViTok,同时保持精细纹理与文本对齐语义,并支持从粗到细的多模态表征策略。此外,采用扩散模型作为图像解码器以增强生成质量并实现高效超分辨率。ILLUME+ 在统一多模态大模型中采用连续输入、离散输出架构,结合渐进式训练,支持视觉分词器、多模态大模型与扩散解码器的动态分辨率。该设计使模型可在多样任务中灵活、高效地进行上下文感知的图像编辑与生成。ILLUME+(3B)在多模态理解、生成与编辑基准测试中表现媲美现有统一模型及专用模型,具备可扩展性与通用性,为未来多模态应用提供坚实基础。

原文摘要 · Abstract (English)

We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation. Existing unified models have struggled to simultaneously handle the three fundamental capabilities in a unified model: understanding, generation, and editing. Models like Chameleon and EMU3 utilize VQGAN for image discretization, due to the lack of deep semantic interaction, they lag behind specialist models like LLaVA in visual understanding tasks. To mitigate this, LaViT and ILLUME employ semantic encoders for tokenization, but they struggle with image editing due to poor texture preservation. Meanwhile, Janus series decouples the input and output image representation, limiting their abilities to seamlessly handle interleaved image-text understanding and generation. In contrast, ILLUME+ introduces a unified dual visual tokenizer, DualViTok, which preserves both fine-grained textures and text-aligned semantics while enabling a coarse-to-fine image representation strategy for multimodal understanding and generation. Additionally, we employ a diffusion model as the image detokenizer for enhanced generation quality and efficient super-resolution. ILLUME+ follows a continuous-input, discrete-output scheme within the unified MLLM and adopts a progressive training procedure that supports dynamic resolution across the vision tokenizer, MLLM, and diffusion decoder. This design allows for flexible and efficient context-aware image editing and generation across diverse tasks. ILLUME+ (3B) exhibits competitive performance against existing unified MLLMs and specialized models across multimodal understanding, generation, and editing benchmarks. With its strong performance, ILLUME+ provides a scalable and versatile foundation for future multimodal applications. Project Page: https://illume-unified-mllm.github.io/.

多模态模型图像生成扩散模型视觉分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。