arXiv:2601.20742cs.CV2026-01被引 2

压缩效率揭示智能本质,统一视觉编码与视觉标记技术

Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification

  • 从压缩效率视角统一视觉编码与视觉标记技术
  • 揭示压缩与模型性能间的内在权衡关系
  • 为多模态大模型、AIGC等应用提供新范式

压缩效率常与模型性能正相关,支撑了人工智能特别是多模态大语言模型(MLLMs)的发展。传统视觉编码基于信息论,已形成广泛应用于多媒体系统的国际标准(如H.264/265)。新兴的生成式多模态大模型中的视觉标记技术,同样以高语义保真度、低计算成本为目标。本文系统梳理两大技术体系,从优化角度统一二者,并揭示其核心机制。基于统一框架,双向赋能技术演进,预测下一代视觉编解码器与标记技术。实验表明,面向任务的标记技术在多模态大模型、AI生成内容及具身智能中具有巨大潜力,未来有望像传统编解码器一样,发展出通用、高效、标准化的智能标记技术。

原文摘要 · Abstract (English)

"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLMs/MLLMs), where compression efficiency often correlates with improved model performance and capabilities. For compression, classical visual coding based on traditional information theory has developed over decades, achieving great success with numerous international industrial standards widely applied in multimedia (e.g., image/video) systems. Except that, the recent emergingvisual token technology of generative multi-modal large models also shares a similar fundamental objective like visual coding: maximizing semantic information fidelity during the representation learning while minimizing computational cost. Therefore, this paper provides a comprehensive overview of two dominant technique families first -- Visual Coding and Vision Token Technology -- then we further unify them from the aspect of optimization, discussing the essence of compression efficiency and model performance trade-off behind. Next, based on the proposed unified formulation bridging visual coding andvisual token technology, we synthesize bidirectional insights of themselves and forecast the next-gen visual codec and token techniques. Last but not least, we experimentally show a large potential of the task-oriented token developments in the more practical tasks like multimodal LLMs (MLLMs), AI-generated content (AIGC), and embodied AI, as well as shedding light on the future possibility of standardizing a general token technology like the traditional codecs (e.g., H.264/265) with high efficiency for a wide range of intelligent tasks in a unified and effective manner.

视觉编码视觉标记多模态统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。