SemHiTok用分层语义码本统一图像编码,兼顾理解与生成。
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
- 基于预训练语义码本构建像素子码本,分离语义与像素处理结构。
- 在LLaVA-v1.5下实现领先图像重建与多模态理解性能。
- 适合需要统一图像表示的多模态大模型研究者使用。
本文提出SemHiTok,一种基于语义引导的分层码本统一图像分词器,可为多模态理解与生成提供一致的离散表示。近期统一图像分词器受到关注,旨在捕捉高层语义特征以支持理解,同时保留低层像素特征以支持生成。以往方法通过联合训练语义蒸馏与像素重建损失来实现,但因理解与生成对特征层次需求不同,难以取得良好平衡。SemHiTok采用新颖的语义引导分层码本设计,在预训练语义码本基础上构建像素子码本,从结构与训练策略上解耦语义与像素信息,使分词器既能捕获像素级细节,又保持对高层语义的理解能力。实验表明,SemHiTok在LLaVA-v1.5设置下达到领先的图像重建与多模态理解表现。进一步构建的统一多模态大模型在多任务中表现出色。大量实验验证分析,该架构实现了更优的权衡。
原文摘要 · Abstract (English)
In this paper, we introduce SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook that provides consistent discrete representations for multimodal understanding and generation. Recently, unified image tokenizers have sparked exploration within the research community, which is designed to capture high-level semantic features for understanding and retaining low-level pixel features for generation. Previous works attempt to train a unified image tokenizer by combining loss for semantic distillation and pixel reconstruction. However, due to the differing levels of features prioritized by multimodal understanding and generation, joint training methods face significant challenges in achieving a good trade-off. SemHiTok addresses this challenge through a novel semantic-guided hierarchical codebook, which builds pixel sub-codebooks on a pretrained semantic codebook. This design decouples the semantic and pixel in terms of structure and training strategy, enabling the tokenizer to capture pixel features while retaining its ability to comprehend high-level semantic information. Our experiments demonstrate that SemHiTok achieves leading performance in image reconstruction and multimodal understanding under the LLaVA-v1.5 setting. Further, we develop a unified MLLM with SemHiTok, which exhibits superior performance across multimodal understanding and generation tasks. Extensive experiments confirm our analysis, showing that our unified image tokenizer architecture achieves a better trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。