提出五类替代卷积的结构化算子,提升图像处理模型对复杂结构的建模能力。
Beyond Convolution: A Taxonomy of Structured Operators for Learning-Based Image Processing
- 按结构分解、自适应加权、基底优化等五类组织新型算子体系
- 在保持等变性的同时增强对非均匀空间依赖的捕捉能力
- 适合需要高阶结构建模的图像生成与分割任务
卷积操作是现代卷积神经网络的基础,因其简洁性、平移等变性和高效实现而广泛应用。然而,其作为固定、线性、局部平均的结构限制了对低秩分解、自适应基表示和非均匀空间依赖等信号结构性质的建模能力。本文系统构建了一套扩展或替代标准卷积的算子分类体系,分为五类:(i) 基于分解的算子,通过奇异值或张量分解分离结构与噪声成分;(ii) 自适应加权算子,根据空间位置或信号内容调节核贡献;(iii) 基底自适应算子,联合优化分析基底与网络权重;(iv) 积分与核算子,将卷积推广至位置依赖和非线性核;(v) 注意力算子,完全放松局部性假设。每类均提供形式定义、与卷积的结构特性对比及适用任务分析。进一步在可线性性、局部性、等变性、计算成本及图像到图像/图像到标签任务适配性等维度进行综合比较,并指出该领域现存挑战与未来方向。
原文摘要 · Abstract (English)
The convolution operator is the fundamental building block of modern convolutional neural networks (CNNs), owing to its simplicity, translational equivariance, and efficient implementation. However, its structure as a fixed, linear, locally-averaging operator limits its ability to capture structured signal properties such as low-rank decompositions, adaptive basis representations, and non-uniform spatial dependencies. This paper presents a systematic taxonomy of operators that extend or replace the standard convolution in learning-based image processing pipelines. We organise the landscape of alternative operators into five families: (i) decomposition-based operators, which separate structural and noise components through singular value or tensor decompositions; (ii) adaptive weighted operators, which modulate kernel contributions as a function of spatial position or signal content; (iii) basis-adaptive operators, which optimise the analysis bases together with the network weights; (iv) integral and kernel operators, which generalise the convolution to position-dependent and non-linear kernels; and (v) attention-based operators, which relax the locality assumption entirely. For each family, we provide a formal definition, a discussion of its structural properties with respect to the convolution, and a critical analysis of the tasks for which the operator is most appropriate. We further provide a comparative analysis of all families across relevant dimensions -- linearity, locality, equivariance, computational cost, and suitability for image-to-image and image-to-label tasks -- and outline the open challenges and future directions of this research area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。