统一视觉语言模型能同时提升理解和生成能力。
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- 用混合数据训练统一模型,理解与生成能力相互促进。
- 模型越训练,理解与生成效果越强,且跨任务泛化更好。
- 生成时学到的知识可反哺理解,尤其在语言模型内部生效。
近期统一视觉语言模型(VLMs)在融合视觉理解与生成能力方面取得进展,其核心假设是:通过在理解与生成任务上混合训练,可实现双向增强。然而该假设尚未被充分验证。本文系统研究了统一VLMs在理解与生成任务间的泛化能力。设计了一个贴近真实场景的数据集,对多种统一架构进行广泛实验与量化评估。主要发现:第一,混合训练的统一模型在不同架构下均表现出理解与生成的相互增益,且效果随数据量增加而提升;第二,多模态输入与输出空间的对齐程度越高,泛化性能越好;第三,生成任务中获得的知识可迁移至理解任务,且该跨任务泛化发生在基础语言模型中,而非仅限模态适配器。结果表明,统一理解与生成对于构建高效VLM至关重要,为模型设计提供关键指导。
原文摘要 · Abstract (English)
Recent advancements in unified vision-language models (VLMs), which integrate both visual understanding and generation capabilities, have attracted significant attention. The underlying hypothesis is that a unified architecture with mixed training on both understanding and generation tasks can enable mutual enhancement between understanding and generation. However, this hypothesis remains underexplored in prior works on unified VLMs. To address this gap, this paper systematically investigates the generalization across understanding and generation tasks in unified VLMs. Specifically, we design a dataset closely aligned with real-world scenarios to facilitate extensive experiments and quantitative evaluations. We evaluate multiple unified VLM architectures to validate our findings. Our key findings are as follows. First, unified VLMs trained with mixed data exhibit mutual benefits in understanding and generation tasks across various architectures, and this mutual benefits can scale up with increased data. Second, better alignment between multimodal input and output spaces will lead to better generalization. Third, the knowledge acquired during generation tasks can transfer to understanding tasks, and this cross-task generalization occurs within the base language model, beyond modality adapters. Our findings underscore the critical necessity of unifying understanding and generation in VLMs, offering valuable insights for the design and optimization of unified VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。