用可解释的结构化模块构建高效神经网络,提升多维信号建模能力
Interpretable and Frugal Learning Systems Employing Multiresolution Pyramids and Volterra Kernels
- 融合小波、伏尔泰拉核与多分辨率分析,构建可微分可解释模块
- 在气象、音频、医学影像任务中实现高精度反演与分割,如脑部MRI自动分析
- 适合关注模型可解释性与高效计算的科研与医疗应用开发者
深度学习广泛用于处理时序、图像及三维医学影像等多维信号,但其表征缺乏显式信号结构且难以审视。本论文提出基于信号理论、数据与任务目标驱动的可解释学习系统,结合多分辨率分析、小波与滤波器组、多速率表示、非线性伏尔泰拉系统与神经计算图。尺度、方向几何、记忆及非线性输入输出交互以可微分算子模块形式表示,中间变量关联核函数、子带、递归与变换域系数,而非仅依赖黑箱特征通道。设计了快速GPU兼容的D维卷积层、多速率采样层、自然域与小波系数域伏尔泰拉核层、有理多项式级联头、稳定性约束的多维IIR滤波器、小波滤波器组与可学习增益的数字剪切波层。这些模块组成特定任务架构,应用于大气、音频、纹理与医学影像的逆建模、分类与分割。在微波辐射反演中,InVeRt利用小波基上的可学习伏尔泰拉核重构垂直温度与湿度剖面;多分辨率滤波器组编码器配伏尔泰拉头用于高效分类;WaveletViT与ShearViT作为子带Transformer块用于WaveNETR与ShearNETR,方向敏感的图像与MRI分割器;MRILong部署训练好的3D T1加权脑部MRI分割检查点,实现缺血性卒中MRI体积的自动分割与纵向分析。
原文摘要 · Abstract (English)
Deep learning models are widely used to process multidimensional signals such as time series, images, and volumetric medical images, but their learned representations often lack explicit signal structure and are difficult to inspect. This thesis develops model-based, signal-theoretic learning systems guided by data and task objectives. It combines multiresolution analysis, wavelets and filter banks, multirate representations, nonlinear Volterra systems, and neural computation graphs. Scale, directional geometry, memory, and nonlinear input-output interactions are represented as differentiable operator modules trainable by backpropagation. The design keeps intermediate variables tied to kernels, subbands, recursions, and transform-domain coefficients rather than only to opaque feature channels. The thesis formulates fast GPU-compatible D-dimensional convolution layers, multirate sampling layers, Volterra-kernel layers in natural and wavelet coefficient domains, rational polynomial cascade heads, stability-constrained multidimensional IIR filters, wavelet banks, and digital shearlet layers with learnable gains. These modules are composed into task-specific architectures for inverse modeling, classification, and segmentation across atmospheric, audio, texture, and medical-imaging problems. In microwave radiometric inversion, InVeRt retrieves vertical temperature and humidity profiles from microwave brightness temperature observations using learnable Volterra kernels in wavelet bases. Multiresolution filter-bank encoders with Volterra heads are used for efficient classification. WaveletViT and ShearViT serve as subband transformer blocks for WaveNETR and ShearNETR, direction-sensitive segmenters for image and MRI segmentation. MRILong deploys trained 3D T1-weighted brain MRI segmenter checkpoints for automatic segmentation and longitudinal analysis of ischemic stroke MRI volumes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。