用50万级视觉代码本提升多模态理解与生成效果
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- 分层代码本设计,冻结主代码本+可训练子代码本协同
- 50万条目代码本实现高利用率与稳定训练
- 适配扩散模型,零成本生成高质量图像
统一多模态大语言模型在联合推进多模态理解与生成方面展现潜力,现有基于代码本的方法或依赖小规模词汇表(约16K条目),缺乏细粒度语义,或盲目扩展导致令牌利用率低、训练不稳定。我们提出UniCode²,一种分层代码本框架,实现大规模、语义对齐且稳定的视觉标记化。通过聚类数百万个SigLIP序列嵌入,构建50万条目的代码本,在保持视觉-语言对齐的同时扩展容量。稳定性通过分层设计保障:冻结的代码本锚定嵌入空间,可训练代码本优化任务特定语义。这种解耦促进高利用率和鲁棒学习。此外,视觉标记与文本语义对齐,支持与预训练扩散解码器无缝集成,仅需少量调整即可实现高质量视觉合成。UniCode²在多个基准测试中表现优异,证明在不牺牲稳定性、语义一致性或模块性的情况下扩展视觉标记空间的可行性。
原文摘要 · Abstract (English)
Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based methods either rely on small vocabularies (~16K entries) that lack fine-grained semantics or naively scale up, resulting in low token utilization and unstable training. We propose UniCode$^2$, a cascaded codebook framework enabling large-scale, semantically aligned, and stable visual tokenization. By clustering millions of SigLIP sequence embeddings, we build a 500K-entry codebook that preserves vision-language alignment while expanding capacity. Stability is ensured via a cascaded design: a frozen codebook anchors the embedding space, and a trainable codebook refines task-specific semantics. This decoupling promotes high utilization and robust learning. Moreover, the alignment of our visual tokens with textual semantics enables seamless integration with pretrained diffusion decoders, supporting high-quality visual synthesis with minimal adaptation. UniCode^2 delivers strong performance across diverse benchmarks, demonstrating the viability of scaling visual token spaces without sacrificing stability, semantics, or modularity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。