统一视觉生成与理解,用连续令牌实现跨任务协同优化。
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
- 采用连续令牌处理图像,离散令牌处理文本,统一建模多模态输入。
- 合理损失权重下,生成与理解性能均达单任务基线水平或更优。
- 使用强预训练大模型和随机生成顺序,显著提升图像生成质量。
我们提出UniFluid,一种统一的自回归框架,用于联合视觉生成与理解,利用连续视觉令牌。该框架处理多模态图像与文本输入,为文本生成离散令牌,为图像生成连续令牌。尽管图像生成与理解任务存在内在权衡,但通过精心设计的训练策略,二者可相互促进。在适当损失平衡权重下,统一模型在两项任务上的表现达到或超过单任务基线。此外,我们证明在训练中使用更强的预训练语言模型及随机顺序生成,对实现高保真图像生成至关重要。基于Gemma模型系列构建的UniFluid,在图像生成与理解任务上均表现出色,具备强泛化能力,适用于图像编辑、视觉描述生成及问答等下游任务。
原文摘要 · Abstract (English)
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating discrete tokens for text and continuous tokens for image. We find though there is an inherent trade-off between the image generation and understanding task, a carefully tuned training recipe enables them to improve each other. By selecting an appropriate loss balance weight, the unified model achieves results comparable to or exceeding those of single-task baselines on both tasks. Furthermore, we demonstrate that employing stronger pre-trained LLMs and random-order generation during training is important to achieve high-fidelity image generation within this unified framework. Built upon the Gemma model series, UniFluid exhibits competitive performance across both image generation and understanding, demonstrating strong transferability to various downstream tasks, including image editing for generation, as well as visual captioning and question answering for understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。