测试视觉语言模型能识别多小的图像细节,发现12像素后感知能力就饱和了。
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

- 设计新基准FineSightBench,用4-48px小图测试模型感知与推理能力
- 12px以下感知能力显著下降,但计数和排序错误在大图中仍普遍存在
- 揭示当前模型在细粒度视觉推理上的根本缺陷,适合关注模型局限性的研究者
近期视觉语言模型(VLMs)在多模态理解与推理方面表现优异,但其细粒度视觉感知能力仍待深入探索。一个自然延伸的问题是:视觉语言模型能可靠识别多小的视觉模式?为此,我们提出FineSightBench,一个系统性探测该极限的新基准,通过在4–48px的可控尺度下,将感知任务(字母、形状、物体的像素级识别)与推理任务(空间关系、计数、顺序判断)分离。对主流模型的全面实验与故障模式分析显示,感知能力在约12px处趋于饱和,而推理能力即使在较大尺度上也持续受限,存在明显的计数错误与序列错误。这些发现揭示了当前视觉语言模型在细尺度视觉推理中的根本缺陷,亟需更严格的评估方法。
原文摘要 · Abstract (English)
Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored. A natural extension of ``How many r are there in Strawberry?'' asks: how small a visual pattern can a VLM reliably perceive? As such, we introduce FineSightBench, a new benchmark that systematically probes this limit by separating perception tasks (pixel-level recognition of letters, shapes, objects) from reasoning tasks (spatial reasoning, counting, ordering over small targets) across controlled scales of 4--48px. Through comprehensive experiments and detailed failure mode analysis on state-of-the-art models, we reveal a sharp dissociation: perception saturates around 12px, while reasoning remains limited even at larger scales, with persistent numeracy and sequence errors. These findings expose fundamental deficiencies in VLMs' fine-scale visual reasoning that demand more rigorous evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。