arXiv:2605.24608cs.AIcs.CV2026-05

用格理论揭示卷积网络的代数本质,解释深度为何带来真实表征力。

Lattice theory and algebraic models for deep convolutional learning based on mathematical morphology

  • 基于格论与数学形态学,统一建模卷积、残差与编解码网络
  • 发现标准CNN层组合非幂等,深度带来真正表征能力提升
  • 提出三种真正的幂等开运算结构,适用于模型设计与分析

我们构建了基于格理论和数学形态学的深度卷积架构(CNN、ResNet、UNet等)的严格代数框架。核心工具是平移不变算子的Matheron-Maragos-Banon-Barrera(MMBB)通用表示理论,系统应用于标准深层网络的每一层。主要发现:标准CNN流程(线性卷积+ReLU+平铺最大池化)是一个交叉格算子——卷积在傅里叶下确界半格中为侵蚀,ReLU为格并闭包,最大池化在逐点极大-加法格中为扩张,其组合并非任何单一形态学开运算。第二发现:逐点格中ReLU的上伴随是全局(非局部)算子,在整体非负函数上恒等,否则为−∞,因此无局部形态学侵蚀可与ReLU构成伴随对。这两项结果共同给出了标准深度CNN具有真正表征力的精确代数原因:组合层非幂等。我们识别并完全刻画了三个真正的幂等开运算层:纯极大-加法形态学层(逐点格)、谱维纳层(傅里叶格)和自对偶形态学层。建立了完整的不动点与收敛理论。该框架还统一了最大池化、步幅卷积与拉普拉斯金字塔,给出激活-池化扩张(APD)分解及其正确伴随。

原文摘要 · Abstract (English)

We develop a rigorous algebraic framework for deep convolutional architectures, CNNs, ResNets, and encoder--decoder networks such as UNet, grounded in lattice theory and mathematical morphology. The central tool is the Matheron--Maragos--Banon--Barrera (MMBB) universal representation theory for translation-invariant operators, which we apply systematically to every layer of a standard deep network. The principal finding is that the standard CNN pipeline (linear convolution~$+$ ReLU~$+$ flat max-pooling) is a cross-lattice operator: the convolution is an erosion in the Fourier inf-semilattice while ReLU is a lattice-join closing and max-pooling is a dilation in the pointwise max-plus lattice, and their composition is a morphological opening in neither. A second finding is that the upper adjoint of ReLU in the pointwise lattice is a global (non-local) operator, the identity on globally non-negative functions and $-\infty$ otherwise, so no local morphological erosion can form an adjunction pair with ReLU. These two results together provide the precise algebraic reason why depth in standard CNNs introduces genuine representational power: the composed layer is not idempotent. Three layer designs that are genuine idempotent openings are identified and fully characterised: the pure max-plus morphological layer (pointwise lattice), the spectral Wiener layer (Fourier lattice), and the self-dual morphological layer. We establish a complete fixed-point and convergence theory. The framework also unifies max-pooling, strided convolution, and the Laplacian pyramid under the Goutsias--Heijmans adjoint pyramid theory, and gives the Activation--Pooling Dilation (APD) factorisation with its correct adjoint.

卷积网络格理论形态学代数建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。