arXiv:2602.06914cs.CV2026-02

揭示视觉任务复杂度如何影响模型对细节信息的保留能力

Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs

  • 构建合成数据集与度量指标,分析视觉冗余与信息压缩关系
  • 发现复杂视觉任务训练可提升模型对细粒度特征的保留能力
  • 为下一代视觉大模型训练提供数据复杂度设计依据

视觉大语言模型(VLLM)的视觉能力长期落后于语言能力,尤其在需要精细视觉信息或空间推理的任务中表现不佳。现有研究常将原因归结为视觉冗余:高层视觉信息被均匀分布到多个标记,而细粒度信息被丢弃。本文深入探究这一假设,提出一个专门用于探测不同视觉特征的合成基准数据集,并设计度量方法以分析冗余与压缩的细微关系。通过在多种复杂视觉任务上微调VLLM,我们发现任务复杂度与视觉压缩存在关联,表明足够比例的高复杂度视觉数据对于改变模型的视觉表征分配方式至关重要,进而提升其在复杂视觉任务上的性能。本研究为下一代VLLM的训练策略提供了重要启示。

原文摘要 · Abstract (English)

Vision capabilities in vision large language models (VLLMs) have consistently lagged behind their linguistic capabilities. In particular, numerous benchmark studies have demonstrated that VLLMs struggle when fine-grained visual information or spatial reasoning is required. However, we do not yet understand exactly why VLLMs struggle so much with these tasks relative to others. Some works have focused on visual redundancy as an explanation, where high-level visual information is uniformly spread across numerous tokens and specific, fine-grained visual information is discarded. In this work, we investigate this premise in greater detail, seeking to better understand exactly how various types of visual information are processed by the model and what types of visual information are discarded. To do so, we introduce a simple synthetic benchmark dataset that is specifically constructed to probe various visual features, along with a set of metrics for measuring visual redundancy, allowing us to better understand the nuances of their relationship. Then, we explore fine-tuning VLLMs on a number of complex visual tasks to better understand how redundancy and compression change based upon the complexity of the data that a model is trained on. We find that there is a connection between task complexity and visual compression, implying that having a sufficient ratio of high complexity visual data is crucial for altering the way that VLLMs distribute their visual representation and consequently improving their performance on complex visual tasks. We hope that this work will provide valuable insights for training the next generation of VLLMs.

视觉模型任务复杂度特征压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。