用关键像素集分析模型决策,发现不同架构专注位置和大小差异。
I Am Big, You Are Little; I Am Right, You Are Wrong
- 通过最小必要像素集衡量模型注意力集中程度
- 发现不同架构的像素集中区域和面积有统计差异,如ConvNext与EVA显著不同
- 误分类图像所需像素集更大,适合研究模型可解释性的人参考
图像分类中的机器学习正快速发展,模型种类繁多,选择合适模型日益重要。尽管分类准确率可统计评估,但对模型工作方式的理解仍有限。为深入理解视觉模型的决策过程,我们提出使用最小必要像素集来衡量模型的‘集中度’:即从模型视角捕捉图像本质所需的最少像素。通过比较不同模型在位置、重叠度和面积上的像素集特征,发现不同架构具有显著不同的集中特性,尤其以ConvNeXt和EVA模型最为突出。此外,误分类样本对应的像素集普遍比正确分类的大,表明模型对错误判断的依赖更广泛。
原文摘要 · Abstract (English)
Machine learning for image classification is an active and rapidly developing field. With the proliferation of classifiers of different sizes and different architectures, the problem of choosing the right model becomes more and more important. While we can assess a model's classification accuracy statistically, our understanding of the way these models work is unfortunately limited. In order to gain insight into the decision-making process of different vision models, we propose using minimal sufficient pixels sets to gauge a model's `concentration': the pixels that capture the essence of an image through the lens of the model. By comparing position, overlap, and size of sets of pixels, we identify that different architectures have statistically different concentration, in both size and position. In particular, ConvNext and EVA models differ markedly from the others. We also identify that images which are misclassified are associated with larger pixels sets than correct classifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。