分析视觉令牌的语言特性,发现其与自然语言既有相似又有本质差异。
Analyzing The Language of Visual Tokens
- 从自然语言视角分析视觉令牌的统计规律
- 视觉令牌符合齐普夫定律,但创新度高导致熵高、压缩差
- 缺乏语法结构,适合研究视觉语言建模的优化方向
随着基于Transformer的视觉语言模型(如LLaVA和Chameleon)兴起,图像块被视作离散令牌,类似于自然语言中的词汇,学习视觉与人类语言之间的联合对齐。然而,人们对这些视觉语言的统计行为了解甚少——它们是否遵循类似的频率分布、语法结构或拓扑特征?本文采用以自然语言为中心的方法,分析离散视觉语言,揭示了显著的相似性与根本差异。我们发现,尽管视觉语言符合齐普夫定律,但更高的令牌创新性带来更大的熵和更低的压缩率,且令牌主要表示物体部件,体现中等粒度。同时,视觉语言缺乏连贯的语法结构,导致困惑度更高、层次组织更弱。最后,尽管视觉模型比其他模型更接近自然语言,但其内部凝聚力仍远低于自然语言。这些发现表明,理解离散视觉语言的统计特性有助于设计更高效的计算机视觉模型。
原文摘要 · Abstract (English)
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representation of images. These models often treat image patches as discrete tokens, analogous to words in natural language, learning joint alignments between visual and human languages. However, little is known about the statistical behavior of these visual languages - whether they follow similar frequency distributions, grammatical structures, or topologies as natural languages. In this paper, we take a natural-language-centric approach to analyzing discrete visual languages and uncover striking similarities and fundamental differences. We demonstrate that, although visual languages adhere to Zipfian distributions, higher token innovation drives greater entropy and lower compression, with tokens predominantly representing object parts, indicating intermediate granularity. We also show that visual languages lack cohesive grammatical structures, leading to higher perplexity and weaker hierarchical organization compared to natural languages. Finally, we demonstrate that, while vision models align more closely with natural languages than other models, this alignment remains significantly weaker than the cohesion found within natural languages. Through these experiments, we demonstrate how understanding the statistical properties of discrete visual languages can inform the design of more effective computer vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。