arXiv:2604.01848cs.CV2026-04

现有视觉语言模型在几何变换下严重失效,暴露空间推理短板

Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

  • 通过旋转、缩放等变换测试,发现模型对空间关系敏感
  • 语义越少时性能下降越明显,跨架构均出现此问题
  • 适合关注多模态模型鲁棒性与几何理解的研究者

本文系统研究了当前先进视觉语言模型(VLMs)在基础几何变换下的本质脆弱性。尽管现代VLMs在识别标准朝向物体和描述复杂场景等语义任务上表现优异,但在更根本层面存在系统性缺陷:缺乏可靠判断物体身份所需的鲁棒空间不变性与等变性。我们通过跨符号草图、自然照片及抽象艺术等多样化视觉领域进行评估,发现当语义内容稀疏时,模型性能急剧下降。该现象在不同架构、模型容量和提示策略中均被观测到。结果揭示了当前VLMs在语义理解与空间推理之间存在系统性差距,凸显未来多模态系统需更强的几何基础。

原文摘要 · Abstract (English)

This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they exhibit systematic failures at a more fundamental level: lack of robust spatial invariance and equivariance required to reliably determine object identity under simple rotations, scaling, and identity transformations. We demonstrate this limitation through a systematic evaluation across diverse visual domains, including symbolic sketches, natural photographs, and abstract art. Performance drops sharply as semantic content becomes sparse, and this behavior is observed across architectures, model capacities, and prompting strategies. Overall, our results reveal a systematic gap between semantic understanding and spatial reasoning in current VLMs, highlighting the need for stronger geometric grounding in future multimodal systems.

视觉语言模型空间推理几何不变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。