研究探针数据对图像分类模型内部概念识别的影响。
On the Performance of Concept Probing: The Influence of the Data (Extended Version)
- 用不同数据训练探针模型,测试其概念识别能力差异。
- 发现探针性能显著受训练数据影响,数据相关性决定效果。
- 开源两个主流数据集的概念标签,方便后续研究。
概念探针近年来受到越来越多关注,用于帮助解释人工神经网络,因其规模庞大且具非符号性,难以直接由人类理解。概念探针通过训练额外分类器,将模型内部表示映射到人类定义的概念,从而实现对神经网络的可解释性观察。现有研究主要聚焦被探针模型或探针本身,较少关注训练探针所需的数据。本文填补这一空白,聚焦图像分类任务中的概念探针,系统研究训练探针所用数据对其性能的影响。同时,我们为两个广泛应用的数据集(CIFAR-10 和 ImageNet)提供了概念标签,以支持未来研究。
原文摘要 · Abstract (English)
Concept probing has recently garnered increasing interest as a way to help interpret artificial neural networks, dealing both with their typically large size and their subsymbolic nature, which ultimately renders them unfeasible for direct human interpretation. Concept probing works by training additional classifiers to map the internal representations of a model into human-defined concepts of interest, thus allowing humans to peek inside artificial neural networks. Research on concept probing has mainly focused on the model being probed or the probing model itself, paying limited attention to the data required to train such probing models. In this paper, we address this gap. Focusing on concept probing in the context of image classification tasks, we investigate the effect of the data used to train probing models on their performance. We also make available concept labels for two widely used datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。