arXiv:2607.25527cs.CVcs.AI2026-07

用极少数据和算力训练出能理解又能生成图像的紧凑模型。

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

论文配图:Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
图 1 · 摘自论文原文
  • 用冻结的视觉编码器+混合视觉标记,兼顾理解与生成需求。
  • 仅用1560万数据、约2000美元成本,性能达顶尖水平。
  • 适合想低成本开发统一多模态模型的研究者使用。

将视觉理解与生成统一于单一模型具有巨大潜力,但因计算与数据需求高,且两类任务所需视觉特征存在冲突,仍面临挑战。为此,我们提出Argus-Unified,一种低计算与数据需求的紧凑高效统一多模态模型。该模型不从零对齐模态,而是有效利用预训练视觉语言模型(VLM)提供的强多模态先验。具体地,引入混合视觉标记:保留连续标记用于理解,同时在冻结的统一视觉编码器基础上学习离散标记用于生成。训练分两阶段进行:第一阶段在冻结视觉编码器上学习量化器与图像解码器;第二阶段使用预训练VLM初始化的LLM进行统一多模态建模。仅使用1560万数据与约2000美元成本,我们证明统一多模态模型可经济高效训练,并在理解任务上达到GQA、POPE、VQAv2等基准的最先进水平,生成质量与专用视觉编码器模型(如Janus、Janus-Pro)相当,成本仅为后者的十分之一,数据量减少五倍。我们期待Argus-Unified成为降低统一模型研发门槛的实用基线。

原文摘要 · Abstract (English)

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified effectively leverages pretrained vision-language models (VLMs) that provide strong multimodal priors. Specifically, we introduce hybrid visual tokens that preserve continuous tokens for understanding while learning discrete tokens for generation from a frozen unified vision encoder. Our training pipeline includes two stages: the first stage learns a quantizer and image decoder on top of the frozen vision encoder, the second stage trains the LLM initialized from a pretrained VLM for the unified multimodal modeling. Using by far the least amount of data (15.6M) and the lowest cost (~$2,000), we demonstrate that unified multimodal models can be trained economically while achieving strong performance in both understanding and generation. Notably, our model attains state-of-the-art multimodal understanding on GQA, POPE, and VQAv2, and competitive generation quality compared to models with dedicated vision encoders (e.g., Janus, Janus-Pro), all at ~10x lower cost and with ~5x less data. We envision Argus-Unified as a useful baseline that lowers the development barrier for unified models.

统一模型图像生成低成本训练多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。