arXiv:2505.15970cs.CVcs.LG2025-05CVPR被引 6

用稀疏自编码器揭示视觉模型如何隐式编码图像分类层级结构。

Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders

  • 通过稀疏自编码器解析视觉模型内部表示,挖掘语义特征。
  • 发现DINOv2模型各层激活中存在隐含的分类层级关系。
  • 适合研究模型表征机制与语义结构对齐的研究者参考。

ImageNet的层级分类体系为分析深度视觉模型学习到的表征提供了有价值的视角。本文利用稀疏自编码器(SAEs)系统分析视觉模型对ImageNet层级结构的编码方式。SAEs在大语言模型中已用于发现语义有意义的特征,本文将其拓展至视觉模型,探究其学习表征是否与ImageNet分类体系的本体结构一致。结果表明,SAEs能揭示模型激活中的层级关系,表现出对分类层级的隐式编码。我们分析了主流视觉基础模型DINOv2在不同层间表示的一致性,发现随着层数加深,类别标记(class token)的信息量逐步增加,表明深层视觉模型逐步内化层级分类信息。本研究建立了一套系统的视觉模型层级分析框架,凸显了SAEs在探测深度网络语义结构方面的潜力。

原文摘要 · Abstract (English)

The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In this work, we conduct a comprehensive analysis of how vision models encode the ImageNet hierarchy, leveraging Sparse Autoencoders (SAEs) to probe their internal representations. SAEs have been widely used as an explanation tool for large language models (LLMs), where they enable the discovery of semantically meaningful features. Here, we extend their use to vision models to investigate whether learned representations align with the ontological structure defined by the ImageNet taxonomy. Our results show that SAEs uncover hierarchical relationships in model activations, revealing an implicit encoding of taxonomic structure. We analyze the consistency of these representations across different layers of the popular vision foundation model DINOv2 and provide insights into how deep vision models internalize hierarchical category information by increasing information in the class token through each layer. Our study establishes a framework for systematic hierarchical analysis of vision model representations and highlights the potential of SAEs as a tool for probing semantic structure in deep networks.

视觉模型层级结构稀疏自编码器表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。