arXiv:2511.20931cs.CVcs.AI2025-11被引 1

用开放词汇模型生成解释,让神经元激活与任意概念对齐。

Open Vocabulary Compositional Explanations for Neuron Alignment

  • 用开放词汇分割掩码定义任意概念,突破预设概念限制。
  • 在多个数据集上实现可量化且更易懂的组合解释。
  • 适合需要灵活探究神经元语义的科研人员使用。

神经元是深度神经网络的基本单元,其连接关系使AI取得前所未有的成果。为理解神经元如何编码信息,组合解释利用概念间的逻辑关系,表达神经元激活与人类知识的空间对齐。然而,现有方法依赖人工标注数据集,仅适用于特定领域和预定义概念。本文提出一种面向视觉领域的框架,允许用户针对任意概念和数据集探测神经元。该框架通过开放词汇语义分割生成掩码,分三步完成:指定任意概念、用开放词汇模型生成分割掩码、从掩码推导组合解释。论文在定量指标和人类可解释性上对比了新方法与以往方法,分析了从人工标注转向模型标注时解释差异,并展示了框架在任务和属性灵活性上的新增能力。

原文摘要 · Abstract (English)

Neurons are the fundamental building blocks of deep neural networks, and their interconnections allow AI to achieve unprecedented results. Motivated by the goal of understanding how neurons encode information, compositional explanations leverage logical relationships between concepts to express the spatial alignment between neuron activations and human knowledge. However, these explanations rely on human-annotated datasets, restricting their applicability to specific domains and predefined concepts. This paper addresses this limitation by introducing a framework for the vision domain that allows users to probe neurons for arbitrary concepts and datasets. Specifically, the framework leverages masks generated by open vocabulary semantic segmentation to compute open vocabulary compositional explanations. The proposed framework consists of three steps: specifying arbitrary concepts, generating semantic segmentation masks using open vocabulary models, and deriving compositional explanations from these masks. The paper compares the proposed framework with previous methods for computing compositional explanations both in terms of quantitative metrics and human interpretability, analyzes the differences in explanations when shifting from human-annotated data to model-annotated data, and showcases the additional capabilities provided by the framework in terms of flexibility of the explanations with respect to the tasks and properties of interest.

神经元解释开放词汇组合解释可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。