提出新型视觉分词方法,兼顾生成与理解性能。
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- 融合连续与离散分词优势,按图像复杂度自适应分配原型数。
- 在重建、检索和分类任务上超越专用连续/离散分词器。
- 适合追求统一模型效率与性能的多模态研究者。
单一多模态大模型中理解与生成的统一仍面临挑战,主要源于连续与离散视觉分词之间的二元对立。连续分词器(CT)虽能连接多个独立训练的理解与生成模块并取得优异性能,但依赖复杂的多阶段流程,工程开销大;离散分词器(DT)通过量化图像为基本单元实现概念上的简洁,却不可避免导致信息损失与性能下降。为此,受光的波粒二象性启发,我们提出连续-离散双态视觉分词器(CDD-VT)。将视觉数据视为来自量化码本的可变组合图像原型,关键在于:根据样本复杂度自适应确定使用原型数量——简单样本用少量原型,模拟离散分词;复杂样本用大量原型,逼近连续分词。设计两大核心组件:多样量化原型(DQP),增强原型正交性以更高效填充信息空间;动态原型分配器(DPA),评估样本复杂度以确定最优原型集合。大量实验表明,CDD-VT在重建、检索和分类任务上均优于专门化的连续与离散分词器,在简洁且可扩展的多模态大模型中实现了更强性能。
原文摘要 · Abstract (English)
The unification of understanding and generation within a single multi-modal large model (MLLM) remains one significant challenge, largely due to the dichotomy between continuous and discrete visual tokenizations. Continuous tokenizer (CT) achieves strong performance by bridging multiple independently-trained understanding modules and generation modules, but suffers from complex multi-stage pipelines and substantial engineering overhead. Conversely, discrete tokenizers (DT) offer a conceptually elegant idea by quantizing each image into a primitive, but inevitably leading to information loss and performance degradation. To resolve this tension, we question the binary choice between CT and DT, inspired by the wave-particle duality of light, and propose the Continuous-Discrete Dualistic Visual Tokenizer (CDD-VT). We treat visual data as a flexible composition of image primitives derived from quantized codebooks, with the crucial insight that the primitive number assigned to each visual sample is adaptively determined according to its complexity: simple instances use a few primitives, emulating discrete tokenization, while complex instances use many, approximating continuous tokenization. Two core components are designed: Diverse Quantitative Primitives, which encourage primitives orthogonality to better populate information space, and Dynamic Primitive Allocator, which assesses sample complexity to determine the optimal set of primitives. Extensive experiments on reconstruction, retrieval and classification show that CDD-VT achieves superior performance over to specialized CT and DT, effectively getting strong result within a concise and scalable MLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。