arXiv:2502.10273cs.CVcs.AI2025-02被引 7

测试大模型在变色、变大小、变形状下的感知稳定性,发现表现差异显著。

Probing Perceptual Constancy in Large Vision-Language Models

  • 用236个实验测试155个视觉语言模型的感知恒常性
  • 形状恒常性表现明显不同于颜色和大小恒常性
  • 结果揭示模型对不同感知变化的敏感度差异

感知恒常性是指在感官输入变化(如距离、角度或光照)时仍能保持对物体稳定感知的能力,这对动态世界中的视觉理解至关重要。本文研究了当前视觉语言模型(VLMs)的该能力。通过在颜色、尺寸和形状三个领域进行236项实验,评估了155个VLMs,涵盖单图与视频形式的经典认知任务及真实场景下的新任务。结果表明,模型在各领域的表现存在显著差异,其中形状恒常性表现与颜色和尺寸恒常性明显分离。

原文摘要 · Abstract (English)

Perceptual constancy is the ability to maintain stable perceptions of objects despite changes in sensory input, such as variations in distance, angle, or lighting. This ability is crucial for visual understanding in a dynamic world. Here, we explored such ability in current Vision Language Models (VLMs). In this study, we evaluated 155 VLMs using 236 experiments across three domains: color, size, and shape constancy. The experiments included single-image and video adaptations of classic cognitive tasks, along with novel tasks in in-the-wild conditions. We found significant variability in VLM performance across these domains, with model performance in shape constancy clearly dissociated from that of color and size constancy.

视觉语言模型感知恒常性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。