一个统一模型同时完成视觉理解与生成,简化结构还效果顶尖。
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

- 用统一自回归框架处理视觉理解和生成,无需扩散模型等额外模块。
- 在多个视觉语言任务上接近当前最佳性能,证明其有效性。
- 适合需要高效多模态生成与理解的工业级应用开发者参考。
VILA-U 是一个整合视频、图像和语言理解与生成的统一基础模型。传统视觉语言模型(VLM)对视觉理解与生成使用独立模块,易导致特征错位且结构复杂。VILA-U 则采用单一自回归下一个标记预测框架完成两项任务,无需额外组件如扩散模型。该方法不仅简化了模型架构,还在视觉语言理解与生成任务中达到近似业界领先水平。其成功主要归因于两点:一是预训练阶段统一视觉塔将离散视觉标记与文本输入对齐,增强视觉感知;二是基于高质量数据集的自回归图像生成可达到与扩散模型相当的质量。因此,VILA-U 能以全标记化自回归框架实现与更复杂模型相媲美的表现。
原文摘要 · Abstract (English)
VILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。