arXiv:2501.12327cs.CV2025-01被引 46

统一视觉理解与生成的自回归多模态大模型,支持图文互动生成。

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

  • 基于自回归框架,统一处理视觉理解与生成任务。
  • 在多个视觉基准上超越LLaVA-1.5,生成质量显著提升。
  • 适合需要图文双向生成与复杂视觉推理的研究者。

我们提出VARGPT,一种新型多模态大语言模型(MLLM),在单一自回归框架中统一视觉理解与生成。VARGPT采用下一标记预测进行视觉理解,下一尺度预测实现视觉自回归生成。该模型创新性地扩展了LLaVA架构,在多模态大模型中实现了高效级联式自回归视觉生成,同时无缝支持混合模态输入输出。通过在精心构建的数据集上进行三阶段统一训练,包括预训练及两个混合视觉指令微调阶段,有效对齐视觉与文本特征,增强指令遵循能力,并提升生成质量。尽管基于LLaVA架构,VARGPT在视觉问答与推理等任务上显著优于LLaVA-1.5。其天然支持自回归视觉生成与指令到图像合成,展现强大的视觉理解与生成双重能力。

原文摘要 · Abstract (English)

We present VARGPT, a novel multimodal large language model (MLLM) that unifies visual understanding and generation within a single autoregressive framework. VARGPT employs a next-token prediction paradigm for visual understanding and a next-scale prediction paradigm for visual autoregressive generation. VARGPT innovatively extends the LLaVA architecture, achieving efficient scale-wise autoregressive visual generation within MLLMs while seamlessly accommodating mixed-modal input and output within a single model framework. Our VARGPT undergoes a three-stage unified training process on specially curated datasets, comprising a pre-training phase and two mixed visual instruction-tuning phases. The unified training strategy are designed to achieve alignment between visual and textual features, enhance instruction following for both understanding and generation, and improve visual generation quality, respectively. Despite its LLAVA-based architecture for multimodel understanding, VARGPT significantly outperforms LLaVA-1.5 across various vision-centric benchmarks, such as visual question-answering and reasoning tasks. Notably, VARGPT naturally supports capabilities in autoregressive visual generation and instruction-to-image synthesis, showcasing its versatility in both visual understanding and generation tasks. Project page is at: \url{https://vargpt-1.github.io/}

多模态视觉生成自回归大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。