高斯导数网络在图像缩放变化下表现稳定,且可解释性强。
Scale generalisation properties of extended scale-covariant and scale-invariant Gaussian derivative networks on image datasets with spatial scaling variations
- 基于高斯导数的网络设计,支持尺度不变与协变特性。
- 在缩放因子4的测试数据上仍保持良好泛化性能,优于传统网络。
- 支持局部定位与可解释性,适合需要透明决策的场景。
本文深入分析了尺度协变与尺度不变高斯导数网络(GaussDerNets)的尺度泛化能力,并提出概念与算法扩展。在新构建的、缩放因子达4的Fashion-MNIST和CIFAR-10重缩放版本上评估,测试数据中的空间缩放未出现在训练中。实验表明,GaussDerNets在新数据集上具备优异的尺度泛化能力;特征响应在尺度上平均池化有时优于以往的尺度最大池化。在最终层使用空间最大池化可实现非中心物体定位,同时保持尺度泛化。训练中引入跨尺度通道的丢弃(scale-channel dropout)正则化,提升性能与泛化。消融研究显示,基于离散高斯核与中心差分算子的离散化方法表现最优或接近最优。可视化激活图与感受野证实其具有出色的可解释性。
原文摘要 · Abstract (English)
This paper presents an in-depth analysis of the scale generalisation properties of the scale-covariant and scale-invariant Gaussian derivative networks, complemented with both conceptual and algorithmic extensions. For this purpose, Gaussian derivative networks (GaussDerNets) are evaluated on new rescaled versions of the Fashion-MNIST and the CIFAR-10 datasets, with spatial scaling variations over a factor of 4 in the testing data, that are not present in the training data. Additionally, evaluations on the previously existing STIR datasets show that the GaussDerNets achieve better scale generalisation than previously reported for these datasets for other types of deep networks. We first experimentally demonstrate that the GaussDerNets have quite good scale generalisation properties on the new datasets, and that average pooling of feature responses over scales may sometimes also lead to better results than the previously used approach of max pooling over scales. Then, we demonstrate that using a spatial max pooling mechanism after the final layer enables localisation of non-centred objects in image domain, with maintained scale generalisation properties. We also show that regularisation during training, by applying dropout across the scale channels, referred to as scale-channel dropout, improves both the performance and the scale generalisation. In additional ablation studies, we demonstrate that discretisations of GaussDerNets, based on the discrete analogue of the Gaussian kernel in combination with central difference operators, perform best or among the best, compared to a set of other discrete approximations of the Gaussian derivative kernels. Finally, by visualising the activation maps and the learned receptive fields, we demonstrate that the GaussDerNets have very good explainability properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。