新基准测试揭示大模型视觉推理能力短板,尤其在细微干扰下表现不佳。
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- 构建多层级视觉推理评测框架,分离感知、规则与组合推理能力。
- 超1.9万张受控图像测试显示,顶尖模型在微小干扰下接近随机猜测。
- 适合关注多模态模型抽象推理能力的研究者与开发者参考。
视觉语言模型(VLMs)在视觉问答和图像描述等任务上取得显著进展,但其是否真正具备视觉推理能力,还是依赖语言先验仍不明确。为此,我们提出VisRes Bench,一个在自然场景下评估视觉推理能力的基准,无需上下文语言监督。通过三个复杂度层级分析模型行为:一级检测在模糊、纹理变化、遮挡、旋转等扰动下的感知补全与全局图像匹配;二级测试单一属性(如颜色、数量、方向)上的规则推理;三级考察需融合多个视觉属性的组合推理。在超过19,000张受控任务图像上,我们发现当前最先进的VLMs在细微感知扰动下表现接近随机,揭示其抽象能力远低于模式识别水平。本文最后讨论了VisRes作为统一框架对推动多模态研究中抽象视觉推理发展的意义。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform visual reasoning as opposed to relying on linguistic priors remains unclear. To address this, we introduce VisRes Bench, a benchmark designed to study visual reasoning in naturalistic settings without contextual language supervision. Analyzing model behavior across three levels of complexity, we uncover clear limitations in perceptual and relational visual reasoning capacities. VisRes isolates distinct reasoning abilities across its levels. Level 1 probes perceptual completion and global image matching under perturbations such as blur, texture changes, occlusion, and rotation; Level 2 tests rule-based inference over a single attribute (e.g., color, count, orientation); and Level 3 targets compositional reasoning that requires integrating multiple visual attributes. Across more than 19,000 controlled task images, we find that state-of-the-art VLMs perform near random under subtle perceptual perturbations, revealing limited abstraction beyond pattern recognition. We conclude by discussing how VisRes provides a unified framework for advancing abstract visual reasoning in multimodal research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。