arXiv:2412.18387cs.AIcs.LG2024-12被引 1

揭示视觉令牌数量与模型能力的非线性关系,为视觉语言模型设计提供理论依据。

Scaling Capability in Token Space: An Analysis of Large Vision Language Model

  • 构建数学框架分析视觉令牌数与序列距离偏差的关系。
  • 发现低令牌数时性能呈亚线性增长,高令牌数时转为线性增长。
  • 适用于研究视觉语言模型架构优化与训练策略的学者。

大型语言模型在参数量和训练数据量上表现出可预测的扩展规律。本研究探究视觉语言模型中视觉令牌数量是否也存在类似扩展关系。通过建立数学框架,刻画视觉令牌数量与视觉引用序列间距离偏差期望值的关系。理论分析揭示两种不同扩展模式:视觉令牌较少时为亚线性扩展,较多时转为线性扩展。这与模型性能关系式 $S(n) \ approx c / n^{α(n)}$ 一致,其中扩展指数与视觉令牌表示间的相关结构有关。在多个视觉语言基准上的实证验证表明,模型性能与扩展关系预测高度吻合。研究成果通过理论框架深化了对变压器中视觉令牌扩展机制的理解,补充了经验观察。

原文摘要 · Abstract (English)

Large language models have demonstrated predictable scaling behaviors with respect to model parameters and training data. This study investigates whether a similar scaling relationship exist for vision-language models with respect to the number of vision tokens. A mathematical framework is developed to characterize a relationship between vision token number and the expected divergence of distance between vision-referencing sequences. The theoretical analysis reveals two distinct scaling regimes: sublinear scaling for less vision tokens and linear scaling for more vision tokens. This aligns with model performance relationships of the form \(S(n) \approx c / n^{α(n)}\), where the scaling exponent relates to the correlation structure between vision token representations. Empirical validations across multiple vision-language benchmarks show that model performance matches the prediction from scaling relationship. The findings contribute to understanding vision token scaling in transformers through a theoretical framework that complements empirical observations.

视觉语言模型扩展规律令牌数量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。