新基准揭示大模型全局视觉感知严重不足
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- 用拓扑结构评估模型全局视觉感知能力
- 所有模型在粗粒度下均接近随机水平
- 越强模型表现越差,暗示规模扩张无效
大型视觉语言模型(LVLM)通常将编码器提取的视觉特征与预训练大语言模型对齐,但这一机制使视觉感知模块成为瓶颈,限制了整体性能。传统评测基准虽富含视觉语义,却存在不可避免的局部捷径,易高估模型感知能力。本文提出TopoPerception,基于拓扑特性严格评估LVLM在多粒度下的全局视觉感知能力。由于拓扑依赖图像全局结构且对局部特征不变,该基准可实现无捷径的全局感知评估,与语义丰富任务本质不同。我们在多个先进模型上测试发现,即使在最粗粒度下,所有模型表现均接近随机水平,表明其严重缺乏全局视觉感知能力。值得注意的是,模型家族内存在一致趋势:推理能力越强的模型,准确率反而越低。这说明单纯扩大模型规模无法解决此缺陷,甚至可能加剧问题。未来改进可能需要新的训练范式或架构。数据与代码已公开于https://github.com/Wenhao-Zhou/TopoPerception。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) typically align visual features from an encoder with a pre-trained Large Language Model (LLM). However, this makes the visual perception module a bottleneck, which constrains the overall capabilities of LVLMs. Conventional evaluation benchmarks, while rich in visual semantics, often contain unavoidable local shortcuts that can lead to an overestimation of models' perceptual abilities. Here, we introduce TopoPerception, a benchmark that leverages topological properties to rigorously evaluate the global visual perception capabilities of LVLMs across various granularities. Since topology depends on the global structure of an image and is invariant to local features, TopoPerception enables a shortcut-free assessment of global perception, fundamentally distinguishing it from semantically rich tasks. We evaluate state-of-the-art models on TopoPerception and find that even at the coarsest perceptual granularity, all models perform no better than random chance, indicating a profound inability to perceive global visual features. Notably, a consistent trend emerge within model families: more powerful models with stronger reasoning capabilities exhibit lower accuracy. This suggests that merely scaling up models is insufficient to address this deficit and may even exacerbate it. Progress may require new training paradigms or architectures. TopoPerception not only exposes a critical bottleneck in current LVLMs but also offers a lens and direction for improving their global visual perception. The data and code are publicly available at: https://github.com/Wenhao-Zhou/TopoPerception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。