arXiv:2412.06639cs.CVcs.AI2024-12NeurIPS被引 10

用概念分解方法揭示视觉Transformer特征空间的深层差异

Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers

  • 将特征空间分解为任意流形表示的概念,实现细粒度对齐分析
  • 发现监督训练越多,模型特征的语义结构越弱,存在显著退化现象
  • 适合研究模型表征机制、对比不同训练方式的科研人员

视觉变换器(ViTs)可采用从全监督到自监督等多种训练范式,不同协议常导致显著不同的特征空间。现有对齐分析仅以单一标量衡量关系,掩盖了共享特征中通用与独特成分的差异。本文结合概念发现与对齐分析,将对齐细化至具体概念层面。提出将概念定义为可能存在的最一般结构——任意流形,通过隐藏特征到流形的距离描述其语义。使用广义兰德指数并进行概念对间划分,实现概念级对齐度量。消融实验验证新定义优于线性基线。对四种不同ViT的分析表明:监督程度越高,学习到的表征语义结构越弱。

原文摘要 · Abstract (English)

Vision transformers (ViTs) can be trained using various learning paradigms, from fully supervised to self-supervised. Diverse training protocols often result in significantly different feature spaces, which are usually compared through alignment analysis. However, current alignment measures quantify this relationship in terms of a single scalar value, obscuring the distinctions between common and unique features in pairs of representations that share the same scalar alignment. We address this limitation by combining alignment analysis with concept discovery, which enables a breakdown of alignment into single concepts encoded in feature space. This fine-grained comparison reveals both universal and unique concepts across different representations, as well as the internal structure of concepts within each of them. Our methodological contributions address two key prerequisites for concept-based alignment: 1) For a description of the representation in terms of concepts that faithfully capture the geometry of the feature space, we define concepts as the most general structure they can possibly form - arbitrary manifolds, allowing hidden features to be described by their proximity to these manifolds. 2) To measure distances between concept proximity scores of two representations, we use a generalized Rand index and partition it for alignment between pairs of concepts. We confirm the superiority of our novel concept definition for alignment analysis over existing linear baselines in a sanity check. The concept-based alignment analysis of representations from four different ViTs reveals that increased supervision correlates with a reduction in the semantic structure of learned representations.

视觉Transformer表征分析概念分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。