统一视觉编码让模型同时懂图像和生成图像,效果领先。
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- 用离散+连续表示融合图像语义与细节
- 在多个基准上达到顶尖性能,超越现有方法
- 适合需要理解与生成一体的多模态研究者
我们提出UniToken,一种自回归生成模型,通过结合离散与连续表示来编码视觉输入,实现统一视觉理解与图像生成任务的无缝集成。与以往依赖单一视觉表示的方法不同,该统一视觉编码框架同时捕捉高层语义与低层细节,提供多维信息,使不同任务可依据自身特性选择性吸收领域知识。深入实验揭示了构建兼具视觉理解与图像生成能力的统一模型的关键原则。在多个主流基准上的广泛评估表明,UniToken表现优异,超越现有方法。这些结果确立了UniToken在该领域的坚实基础。代码与模型已公开于https://github.com/SxJyJay/UniToken。
原文摘要 · Abstract (English)
We introduce UniToken, an auto-regressive generation model that encodes visual inputs through a combination of discrete and continuous representations, enabling seamless integration of unified visual understanding and image generation tasks. Unlike previous approaches that rely on unilateral visual representations, our unified visual encoding framework captures both high-level semantics and low-level details, delivering multidimensional information that empowers heterogeneous tasks to selectively assimilate domain-specific knowledge based on their inherent characteristics. Through in-depth experiments, we uncover key principles for developing a unified model capable of both visual understanding and image generation. Extensive evaluations across a diverse range of prominent benchmarks demonstrate that UniToken achieves state-of-the-art performance, surpassing existing approaches. These results establish UniToken as a robust foundation for future research in this domain. The code and models are available at https://github.com/SxJyJay/UniToken.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。