arXiv:2602.02539cs.LGcs.CV2026-02被引 1

揭示视觉令牌的信息承载极限,给出通用压缩效率指南。

How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs

  • 通过逐步增加图像字符数测试视觉编码器容量,发现三阶段转变现象。
  • 提出概率标度律,统一负载与密度为潜在难度指标,跨模型验证有效。
  • 适用于优化视觉上下文压缩中的精度与效率平衡,尤其适合长文本识别场景。

近期以视觉为中心的方法在长上下文建模方面取得显著进展。以 DeepSeek-OCR 为代表,这些模型将渲染文本编码为连续视觉标记,实现高压缩率且不损失识别精度。然而,将视觉编码器视为具有有限表征能力的有损信道,引出一个根本性问题:视觉标记的信息上限是多少?为探究这一极限,我们通过逐步增加图像内信息量(字符数)进行受控压力测试。观察到明显的相变现象,包含三个阶段:近乎完美的稳定阶段、误差方差上升的不稳定性阶段,以及彻底崩溃阶段。我们分析了这些转变的机制成因,并识别关键影响因素。此外,我们建立了一个概率标度律,将平均视觉标记负载与视觉密度统一为潜在难度指标。在多种视觉-语言模型上的广泛实验表明该标度律具有普适性,为视觉上下文压缩中的效率-精度权衡提供了关键实证指导。

原文摘要 · Abstract (English)

Recent vision-centric approaches have made significant strides in long-context modeling. Represented by DeepSeek-OCR, these models encode rendered text into continuous vision tokens, achieving high compression rates without sacrificing recognition precision. However, viewing the vision encoder as a lossy channel with finite representational capacity raises a fundamental question: what is the information upper bound of visual tokens? To investigate this limit, we conduct controlled stress tests by progressively increasing the information quantity (character count) within an image. We observe a distinct phase-transition phenomenon characterized by three regimes: a near-perfect Stable Phase, an Instability Phase marked by increased error variance, and a total Collapse Phase. We analyze the mechanical origins of these transitions and identify key factors. Furthermore, we formulate a probabilistic scaling law that unifies average vision token load and visual density into a latent difficulty metric. Extensive experiments across various Vision-Language Models demonstrate the universality of this scaling law, providing critical empirical guidance for optimizing the efficiency-accuracy trade-off in visual context compression.

视觉编码信息极限标度律多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。