arXiv:2502.05178cs.CV2025-02被引 44

用量化视觉编码统一图文理解与生成,效果更优且无需额外训练。

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

  • 基于二值球面量化构建自编码器,兼顾图像重建与图文对齐。
  • 单模型实现零样本图文理解与文本控制图像生成,性能相当或更优。
  • 支持端到端统一建模,适合需要多任务融合的视觉语言研究者。

我们提出量化图文预训练(QLIP),一种结合先进重建质量与先进零样本图像理解能力的视觉标记化方法。QLIP 采用基于二值球面量化的自编码器,同时优化重建和图文对齐目标。我们首次证明这两个目标无需冲突。通过动态平衡损失项,并采用两阶段训练流程,有效融合大规模图文预训练的数据需求与重建目标带来的内存瓶颈。我们在单个模型上验证了 QLIP 在多模态理解与文本条件图像生成中的有效性。具体而言,QLIP 可作为 LLaVA 的视觉编码器或 LlamaGen 的图像标记器,表现相当甚至更优。最后,我们展示了 QLIP 能够支持统一的混合模态自回归模型,实现理解和生成一体化。

原文摘要 · Abstract (English)

We introduce Quantized Language-Image Pretraining (QLIP), a visual tokenization method that combines state-of-the-art reconstruction quality with state-of-the-art zero-shot image understanding. QLIP trains a binary-spherical-quantization-based autoencoder with reconstruction and language-image alignment objectives. We are the first to show that the two objectives do not need to be at odds. We balance the two loss terms dynamically during training and show that a two-stage training pipeline effectively mixes the large-batch requirements of image-language pre-training with the memory bottleneck imposed by the reconstruction objective. We validate the effectiveness of QLIP for multimodal understanding and text-conditioned image generation with a single model. Specifically, QLIP serves as a drop-in replacement for the visual encoder for LLaVA and the image tokenizer for LlamaGen with comparable or even better performance. Finally, we demonstrate that QLIP enables a unified mixed-modality auto-regressive model for understanding and generation.

视觉编码图文生成统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。