通过分析视觉语言模型的表征几何,揭示其多物体识别失败的内在机制。
The Geometry of Representational Failures in Vision Language Models
- 用概念向量捕捉视觉概念的潜在方向,实现对模型行为的精准操控
- 概念向量几何重叠度与特定错误模式高度相关,可量化解释模型失误
- 适用于研究模型幻觉、识别偏差等认知类缺陷,适合系统性分析者
视觉语言模型(VLMs)在多物体视觉任务中表现出令人困惑的失败,如虚构不存在元素或在干扰中无法识别最相似对象。这些错误虽类似人类的认知限制(如‘绑定问题’),但其在人工系统中的内在机制仍不清晰。本文通过对开放权重的VLMs(Qwen、InternVL、Gemma)进行表征几何分析,提出一种方法以提炼“概念向量”——编码视觉概念的潜在方向。通过控制干预验证,这些向量能可靠地操纵模型在简化和自然场景下的行为(例如强制模型将红花视为蓝色)。我们发现,这些向量间的几何重叠度与特定错误模式显著相关,为理解内部表示如何塑造模型行为并引发视觉失败提供了可量化的框架。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing to identify the most similar objects among distractions. While these errors mirror human cognitive constraints, such as the 'Binding Problem', the internal mechanisms driving them in artificial systems remain poorly understood. Here, we propose a mechanistic insight by analyzing the representational geometry of open-weight VLMs (Qwen, InternVL, Gemma), comparing methodologies to distill "concept vectors'' - latent directions encoding visual concepts. We validate our concept vectors via steering interventions that reliably manipulate model behavior in both simplified and naturalistic vision tasks (e.g., forcing the model to perceive a red flower as blue). We observe that the geometric overlap between these vectors strongly correlates with specific error patterns, offering a grounded quantitative framework to understand how internal representations shape model behavior and drive visual failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。