统一生成文本图像,性能提升13.3%
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
- 将图文转为离散符号,用同一Transformer自回归生成
- 通过渐进式词汇学习,视觉词元逐步融入训练
- 在多任务中表现优异,适合跨模态统一建模
我们提出UGen,一种统一的自回归多模态模型,在文本处理、图像理解与图像生成任务上均表现强劲。UGen将文本和图像转化为离散标记序列,并使用单一Transformer以自回归方式统一生成。为应对统一多模态学习的挑战,UGen采用新颖的渐进式词汇学习机制:视觉标记ID逐步激活并融入训练过程,显著提升统一多模态学习效果。在多个文本与图像任务上的实验表明,相较于基础统一自回归方法,UGen整体性能提升13.3%,且在各项任务中均达到与专用模型相当的竞争力。
原文摘要 · Abstract (English)
We introduce UGen, a unified autoregressive multimodal model that demonstrates strong performance across text processing, image understanding, and image generation tasks simultaneously. UGen converts both texts and images into discrete token sequences and utilizes a single transformer to generate them uniformly in an autoregressive manner. To address the challenges associated with unified multimodal learning, UGen is trained using a novel mechanism, namely progressive vocabulary learning. In this process, visual token IDs are incrementally activated and integrated into the training phase, ultimately enhancing the effectiveness of unified multimodal learning. Experiments on comprehensive text and image tasks show that UGen achieves a significant overall performance improvement of 13.3% compared to the vanilla unified autoregressive method, and it also delivers competitive results across all tasks against several task-specific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。