arXiv:2510.08638cs.CVcs.AI2025-10中稿 · ICLR被引 24

揭示DINOv2如何用概念组合理解视觉,提出新几何解释框架

Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry

  • 用SAE构建3.2万个可解释概念词典,分析任务如何调用不同概念
  • 发现模型表示非严格稀疏,而是由原型凸组合构成的低维局部结构
  • 提出闵可夫斯基表征假说,适合研究视觉模型内部机制的读者

DINOv2常用于识别物体、场景和动作,但其感知本质仍不明。我们以线性表征假说为基础,利用SAE构建了一个包含32,000个单元的概念词典,作为研究核心。第一部分分析不同下游任务如何从词典中提取概念:分类依赖“其他处”概念(在目标对象外处处激活),实现学习到的否定;分割依赖形成连贯子空间的边界检测器;深度估计则利用三种匹配视觉神经科学原理的单目深度线索。第二部分分析SAE学习到概念的几何与统计特性:表示部分稠密而非严格稀疏,词典趋向更高一致性并偏离最大正交理想(格拉斯曼框架);图像内标记占据低维、局部连接集合,移除位置后仍保持稳定。这些迹象表明表示组织超越线性稀疏性。综合观察,我们提出新观点:标记由原型的凸组合形成(如兔子之于动物,棕色之于颜色,蓬松之于纹理)。该结构基于Gardenfors概念空间理论,且符合多头注意力产生凸组合和原型界定区域的机制。我们提出闵可夫斯基表征假说(MRH),并检验其经验特征及其对视觉变换器表示的解释意义。

原文摘要 · Abstract (English)

DINOv2 is routinely deployed to recognize objects, scenes, and actions; yet the nature of what it perceives remains unknown. As a working baseline, we adopt the Linear Representation Hypothesis (LRH) and operationalize it using SAEs, producing a 32,000-unit dictionary that serves as the interpretability backbone of our study, which unfolds in three parts. In the first part, we analyze how different downstream tasks recruit concepts from our learned dictionary, revealing functional specialization: classification exploits "Elsewhere" concepts that fire everywhere except on target objects, implementing learned negations; segmentation relies on boundary detectors forming coherent subspaces; depth estimation draws on three distinct monocular depth cues matching visual neuroscience principles. Following these functional results, we analyze the geometry and statistics of the concepts learned by the SAE. We found that representations are partly dense rather than strictly sparse. The dictionary evolves toward greater coherence and departs from maximally orthogonal ideals (Grassmannian frames). Within an image, tokens occupy a low dimensional, locally connected set persisting after removing position. These signs suggest representations are organized beyond linear sparsity alone. Synthesizing these observations, we propose a refined view: tokens are formed by combining convex mixtures of archetypes (e.g., a rabbit among animals, brown among colors, fluffy among textures). This structure is grounded in Gardenfors' conceptual spaces and in the model's mechanism as multi-head attention produces sums of convex mixtures, defining regions bounded by archetypes. We introduce the Minkowski Representation Hypothesis (MRH) and examine its empirical signatures and implications for interpreting vision-transformer representations.

视觉模型概念解译表示几何注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。