arXiv:2503.23062cs.CV2025-03被引 5

评测大模型对形状纹理的识别能力,发现其低层视觉感知仍有重大缺陷。

Shape and Texture Recognition in Large Vision-Language Models

  • 构建超大规模无监督提取的形状纹理数据集LAS&T
  • 大模型在3D材质识别接近人类,但2D抽象纹理识别严重不足
  • 揭示主流大模型依赖语义特征,缺乏低层视觉建模能力

形状与纹理是视觉感知的基本单元。能否在不依赖朝向、纹理或上下文的情况下识别形状,以及能否独立于物体识别纹理与材质,对通用视觉理解至关重要。本文提出大型形状与纹理数据集(LAS&T),通过从自然图像中无监督提取模式构建,包含超过70万张用于2D/3D形状与纹理识别和检索的图像。该数据集用于评估主流大视觉语言模型(LVLM)在识别与表征2D/3D场景中形状、纹理和材料方面的表现。形状识别测试显示,当图像存在方向、纹理、颜色或环境差异时,大模型性能远低于人类,尤其在多重变换下表现更差;它们主要依赖高层语义特征,难以处理无类别关联的抽象形状。纹理与材料识别方面,领先模型在3D场景中接近人类水平,但在识别简单抽象的2D纹理和形状时显著落后。这一结果在GPT/Gemini/LLama/Qwen等多种主流模型及DINO/CLIP等基础视觉模型中一致出现,暴露了当前视觉语言模型在提取低层视觉特征上的系统性缺陷。相比之下,人类及专为该任务训练的简单网络可达到高准确率。LAS&T数据集已公开可用。

原文摘要 · Abstract (English)

Shapes and textures are the basic building blocks of visual perception. The ability to identify shapes regardless of orientation, texture, or context, and to recognize textures and materials independently of their associated objects, is essential for a general visual understanding of the world. This work introduces the Large Shapes and Textures dataset (LAS&T), a giant collection of highly diverse shapes and textures, created by unsupervised extraction of patterns from natural images. This dataset is used to benchmark how effectively leading Large Vision-Language Models (LVLM/VLM) recognize and represent shapes, textures, and materials in 2D and 3D scenes. For shape recognition, we test the models' ability to match images of identical shapes that differ in orientation, texture, color, or environment. Our results show that the shape-recognition capabilities of LVLMs remain well below human performance, especially when multiple transformations are applied. LVLMs rely predominantly on high-level and semantic features and struggle with abstract shapes lacking class associations. For texture and material recognition, we evaluated the models' ability to identify images with identical textures and materials across different objects and environments. Interestingly, leading LVLMs approach human-level performance in recognizing materials in 3D scenes, yet substantially underperform humans when identifying simpler, more abstract 2D textures and shapes. These results are consistent across a wide range of leading LVLMs (GPT/Gemini/LLama/Qwen) and foundation vision models (DINO/CLIP), exposing major deficiencies in the ability of VLMs to extract low-level visual features. In contrast, humans and simple nets trained directly for these tasks achieve high accuracy. The LAS&T dataset, featuring over 700,000 images for 2D/3D shape and textures recognition and retrieval, is freely available.

视觉理解大模型评测低层视觉纹理识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。