arXiv:2508.13460cs.CV2025-08被引 4

用传统图像编码思想重新审视多模态大模型的分词技术。

Revisiting MLLM Token Technology through the Lens of Classical Visual Coding

  • 建立分词与图像编码的统一框架,实现模块化对比分析。
  • 发现图像编码原理可提升多模态模型的效率与鲁棒性。
  • 为下一代语义图像编解码器提供新设计思路,适合相关研究者。

经典图像编码与多模态大语言模型(MLLM)分词技术的核心目标一致:在保持信息保真度的同时最小化计算成本。本文从长期发展的图像编码领域出发,重新审视MLLM分词技术,包括分词、分词压缩和分词推理。我们提出:(1) 构建统一形式化框架,连接分词技术与图像编码,实现系统性的模块化对比分析;(2) 双向融合洞察,探索图像编码原则如何提升分词技术的效率与鲁棒性,反之亦然,分词范式如何启发下一代语义图像编解码器的设计;(3) 展望有前景的研究方向及关键未解决问题。本研究首次系统性地比较了MLLM分词与图像编码技术,为更高效的多模态模型与更强的视觉编解码器发展铺平道路。

原文摘要 · Abstract (English)

Classical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM token technology, including tokenization, token compression, and token reasoning, through the established principles of long-developed visual coding area. From this perspective, we (1) establish a unified formulation bridging token technology and visual coding, enabling a systematic, module-by-module comparative analysis; (2) synthesize bidirectional insights, exploring how visual coding principles can enhance MLLM token techniques' efficiency and robustness, and conversely, how token technology paradigms can inform the design of next-generation semantic visual codecs; (3) prospect for promising future research directions and critical unsolved challenges. In summary, this study presents the first comprehensive and structured technology comparison of MLLM token and visual coding, paving the way for more efficient multimodal models and more powerful visual codecs simultaneously.

多模态分词技术图像编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。