用尺度空间理论建模深度可分离网络的8类核心滤波器,发现其可被离散高斯差分算子近似。
Modelling and analysis of the 8 filters from the "master key filters hypothesis" for depthwise-separable deep networks in relation to idealized receptive fields based on scale-space theory
- 基于聚类提取8类核心滤波器,用离散高斯差分算子建模其感受野。
- 模型在空间方差匹配或最小化l1/l2误差下,与真实滤波器高度相似。
- 成果适用于高效网络设计与滤波器可解释性研究。
本文分析并建模了从基于ConvNeXt架构的深度可分离网络中通过聚类提取的8个'主键滤波器'。首先计算滤波器绝对值的加权均值与加权方差,支持两个假设:(i) 学习到的滤波器可在空间域上由可分离滤波操作建模;(ii) 非中心滤波器的空间偏移接近半个网格单位。随后,将聚类后的'主键滤波器'建模为离散高斯核平滑后应用差分算子的结果。该建模分两种方式:(i) 每个滤波器使用不同坐标方向的尺度参数;(ii) 所有方向使用相同尺度参数。模型拟合通过强制空间方差相等或最小化离散l1/l2范数实现。实验表明,理想化模型能良好预测真实滤波器,证明深度可分离网络中的学习滤波器可被离散尺度空间滤波器有效近似。
原文摘要 · Abstract (English)
This paper presents the results of analysing and modelling a set of 8 ``master key filters'', which have been extracted by applying a clustering approach to the receptive fields learned in depthwise-separable deep networks based on the ConvNeXt architecture. For this purpose, we first compute spatial spread measures in terms of weighted mean values and weighted variances of the absolute values of the learned filters, which support the working hypotheses that: (i) the learned filters can be modelled by separable filtering operations over the spatial domain, and that (ii) the spatial offsets of the those learned filters that are non-centered are rather close to half a grid unit. Then, we model the clustered ``master key filters'' in terms of difference operators applied to a spatial smoothing operation in terms of the discrete analogue of the Gaussian kernel, and demonstrate that the resulting idealized models of the receptive fields show good qualitative similarity to the learned filters. This modelling is performed in two different ways: (i) using possibly different values of the scale parameters in the coordinate directions for each filter, and (ii) using the same value of the scale parameter in both coordinate directions. Then, we perform the actual model fitting by either (i) requiring spatial spread measures in terms of spatial variances of the absolute values of the receptive fields to be equal, or (ii) minimizing the discrete $l_1$- or $l_2$-norms between the idealized receptive field models and the learned filters. Complementary experimental results then demonstrate the idealized models of receptive fields have good predictive properties for replacing the learned filters by idealized filters in depthwise-separable deep networks, thus showing that the learned filters in depthwise-separable deep networks can be well approximated by discrete scale-space filters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。