自动找模型中适合探测人类概念的层,提升探针有效性
Concept Probing: Where to Find Human-Defined Concepts (Extended Version)
- 根据表征对概念的丰富度和规律性自动选择探测层
- 在多个模型和数据集上验证了方法的有效性
- 适合关注模型内部表征可解释性的研究者
概念探针近年来成为人类窥探人工神经网络内部编码信息的重要手段。通过训练额外分类器,将模型内部表征映射到人类定义的概念。然而探针性能高度依赖所探测的表征层,因此确定合适探测层至关重要。本文提出一种自动识别最佳探测层的方法,依据表征对目标概念的信息量和规律性。我们在多种神经网络模型和数据集上进行了全面实证分析,验证了该方法的可靠性。
原文摘要 · Abstract (English)
Concept probing has recently gained popularity as a way for humans to peek into what is encoded within artificial neural networks. In concept probing, additional classifiers are trained to map the internal representations of a model into human-defined concepts of interest. However, the performance of these probes is highly dependent on the internal representations they probe from, making identifying the appropriate layer to probe an essential task. In this paper, we propose a method to automatically identify which layer's representations in a neural network model should be considered when probing for a given human-defined concept of interest, based on how informative and regular the representations are with respect to the concept. We validate our findings through an exhaustive empirical analysis over different neural network models and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。