用轻量设计实现图像理解与生成的统一,提升效率与效果。
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
- 采用视觉专家+分块令牌机制,简化模型结构
- 在多个数据集上超越同类模型,参数更少性能更强
- 适合追求高效统一多模态系统的研究者使用
大型语言模型(LLMs)的成功已延伸至多模态领域,在图像理解和生成任务中表现卓越。近期致力于构建集成两类能力的统一多模态大模型(MLLMs)的研究展现出良好前景。然而,现有方法常伴随复杂的模型架构或训练流程,增加了训练和扩展难度。本文提出SynerGen-VL,一种无需编码器的高效统一多模态模型,兼具图像理解与生成能力。为解决现有无编码器统一模型的挑战,我们引入令牌折叠机制和基于视觉专家的渐进对齐预训练策略,有效支持高分辨率图像理解的同时降低训练复杂度。在大规模图文混合数据上以统一的下一步词预测目标进行训练后,SynerGen-VL在性能上达到或超过现有无编码器统一模型,且参数规模相当或更小,同时缩小了与专用任务顶尖模型的差距,展示了未来统一多模态模型的可行路径。代码与模型将公开发布。
原文摘要 · Abstract (English)
The remarkable success of Large Language Models (LLMs) has extended to the multimodal domain, achieving outstanding performance in image understanding and generation. Recent efforts to develop unified Multimodal Large Language Models (MLLMs) that integrate these capabilities have shown promising results. However, existing approaches often involve complex designs in model architecture or training pipeline, increasing the difficulty of model training and scaling. In this paper, we propose SynerGen-VL, a simple yet powerful encoder-free MLLM capable of both image understanding and generation. To address challenges identified in existing encoder-free unified MLLMs, we introduce the token folding mechanism and the vision-expert-based progressive alignment pretraining strategy, which effectively support high-resolution image understanding while reducing training complexity. After being trained on large-scale mixed image-text data with a unified next-token prediction objective, SynerGen-VL achieves or surpasses the performance of existing encoder-free unified MLLMs with comparable or smaller parameter sizes, and narrows the gap with task-specific state-of-the-art models, highlighting a promising path toward future unified MLLMs. Our code and models shall be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。