用信息论优化视觉标记压缩,让多模态大模型更高效地理解与生成图像。
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
- 基于信息瓶颈原理,约束图像到标记的信息流动。
- 在3个主流模型上提升理解和生成性能,无需额外训练数据。
- 适合研究多模态模型压缩与高效表示的学者参考。
统一多模态大语言模型(MLLMs)旨在单一框架内整合图像理解与生成能力,其中共享视觉标记器作为唯一接口,将高维图像映射到有限的标记预算中以支持下游多模态推理与合成。然而,现有共享标记设计多依赖架构经验,缺乏明确标准来决定应保留何种信息以同时支持语义抽象与视觉细节。本文从容量受限视角出发,将共享标记器视为计算受限的学习者,其有限表征预算应优先保留可复用结构,而非难利用的高熵变化与冗余。受此启发,我们提出 extbf{ extit{InfoTok}},一种基于信息瓶颈(IB)原理的信息正则化标记机制。InfoTok通过施加互信息(MI)约束,显式控制从图像到共享标记及多模态输出的信息流,实现压缩与任务相关性之间的合理权衡,同时促进跨模态一致性。由于高维视觉表征下互信息难以计算,我们采用可微的近似估计器,包括变分IB形式和基于希尔伯特-施密特独立性准则(HSIC)的替代方案。在三个代表性统一MLLM中集成,不引入额外训练数据,InfoTok一致提升图像理解与生成性能,验证了信息正则化视觉标记化在统一MLLM中的有效性。
原文摘要 · Abstract (English)
Unified multimodal large language models (MLLMs) aim to unify image understanding and image generation within a single framework, where a shared visual tokenizer serves as the sole interface that maps high-dimensional images into a limited token budget for downstream multimodal reasoning and synthesis. However, existing shared-token designs are largely architecture-driven and lack an explicit criterion for what information should be preserved to simultaneously support semantic abstraction and visual detail. In this paper, we adopt a capacity-constrained perspective, viewing the shared tokenizer as a compute-bounded learner whose finite representational budget should prioritize reusable structure over hard-to-exploit high-entropy variations and redundancy. Motivated by this view, we propose \textbf{\textit{InfoTok}}, an information-regularized tokenization mechanism grounded in the Information Bottleneck (IB) principle. InfoTok explicitly controls information flow from images to shared tokens to multimodal outputs by imposing mutual-information (MI) constraints that enforce a principled trade-off between compression and task relevance, while also encouraging cross-modal consistency. Because MI is intractable for high-dimensional visual representations, we instantiate InfoTok with practical, differentiable dependence estimators, including a variational IB formulation and a Hilbert Schmidt Independence Criterion (HSIC) based alternative. Integrated into three representative unified MLLMs without introducing any additional training data, InfoTok consistently improves both image understanding and generation performance. These results support information-regularized visual tokenization as a sound basis for token learning in unified MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。