用数学理论解释CNN如何解决图像逆问题,让黑箱变透明。
An analytic theory of convolutional neural network inverse problems solvers
- 基于最小均方误差框架,结合平移等变性和局部感受野约束,推导出可解释公式。
- 在去噪、补全、去卷积任务中,理论预测与实际输出的PSNR差异小于25dB。
- 适合关注模型可解释性、逆问题理论分析的研究者。
监督式卷积神经网络(CNN)在众多成像逆问题中表现优异,但其理论基础薄弱,常被视为黑箱。本文从最小均方误差(MMSE)估计器出发,引入捕捉CNN两大归纳偏置(平移等变性与有限感受野局部性)的函数约束,推导出一种可解析、可解释且可计算的约束变体——局部等变MMSE(LE-MMSE)。在多种逆问题(去噪、补全、去卷积)、数据集(FFHQ、CIFAR-10、FashionMNIST)和架构(U-Net、ResNet、PatchMLP)上进行大量实验,结果表明该理论预测与神经网络输出高度一致(PSNR ≥ 25dB)。研究还揭示了物理感知与非物理感知估计器的差异、训练分布中高密度区域的影响,以及数据集大小、图像块尺寸等因素的作用。
原文摘要 · Abstract (English)
Supervised convolutional neural networks (CNNs) are widely used to solve imaging inverse problems, achieving state-of-the-art performance in numerous applications. However, despite their empirical success, these methods are poorly understood from a theoretical perspective and often treated as black boxes. To bridge this gap, we analyze trained neural networks through the lens of the Minimum Mean Square Error (MMSE) estimator, incorporating functional constraints that capture two fundamental inductive biases of CNNs: translation equivariance and locality via finite receptive fields. Under the empirical training distribution, we derive an analytic, interpretable, and tractable formula for this constrained variant, termed Local-Equivariant MMSE (LE-MMSE). Through extensive numerical experiments across various inverse problems (denoising, inpainting, deconvolution), datasets (FFHQ, CIFAR-10, FashionMNIST), and architectures (U-Net, ResNet, PatchMLP), we demonstrate that our theory matches the neural networks outputs (PSNR $\gtrsim25$dB). Furthermore, we provide insights into the differences between \emph{physics-aware} and \emph{physics-agnostic} estimators, the impact of high-density regions in the training (patch) distribution, and the influence of other factors (dataset size, patch size, etc).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。