arXiv:2603.14610cs.CVeess.IV2026-03中稿 · the IEEE/CVF Confe…

揭示分类器隐含语义不变性,让模型决策过程可解释。

Make it SING: Analyzing Semantic Invariants in Classifiers

  • 通过特征映射到多模态模型,生成语义等价图像并解释变化
  • ResNet50的不变空间泄露语义信息,而DinoViT更保持语义一致性
  • 适用于单图局部分析或类别级统计,适合模型可解释性研究

所有分类器,包括最先进的视觉模型,都存在由线性映射几何结构部分决定的不变性。这些不变性存在于分类器的零空间中,导致不同输入映射到相同输出。现有方法难以提供人类可理解的语义信息。为此,我们提出语义不变性解析(SING),通过将网络特征映射到多模态视觉语言模型,构建与网络等价的图像,并赋予其语义解释。该方法可生成自然语言描述和视觉示例,揭示隐含的语义变化。SING既可用于单张图像的局部分析,也可用于图像集合的类级与模型级统计分析。实验表明,ResNet50的零空间会泄露相关语义属性,而采用自监督DINO预训练的DinoViT在不变空间中更有效地保持了类别语义一致性。

原文摘要 · Abstract (English)

All classifiers, including state-of-the-art vision models, possess invariants, partially rooted in the geometry of their linear mappings. These invariants, which reside in the null-space of the classifier, induce equivalent sets of inputs that map to identical outputs. The semantic content of these invariants remains vague, as existing approaches struggle to provide human-interpretable information. To address this gap, we present Semantic Interpretation of the Null-space Geometry (SING), a method that constructs equivalent images, with respect to the network, and assigns semantic interpretations to the available variations. We use a mapping from network features to multi-modal vision language models. This allows us to obtain natural language descriptions and visual examples of the induced semantic shifts. SING can be applied to a single image, uncovering local invariants, or to sets of images, allowing a breadth of statistical analysis at the class and model levels. For example, our method reveals that ResNet50 leaks relevant semantic attributes to the null space, whereas DinoViT, a ViT pretrained with self-supervised DINO, is superior in maintaining class semantics across the invariant space.

模型解释不变性视觉语言模型深度学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。