将数学原理与深度学习结合,实现高精度可解释的统计保证模型。
Wahkon: A Statistically Principled Deep RKHS Superposition Network

- 基于核空间正则化与叠加定理构建深层网络,支持逐层复杂度控制。
- 理论证明其估计为最大后验概率解,收敛速度达到最优。
- 在基因数据等场景表现优于传统神经网络,适合追求可信预测的领域。
深度学习擅长预测但常缺乏有限样本保证和校准的不确定性;基于再生核希尔伯特空间(RKHS)的方法虽能提供这些保证,却难以适应高维数据。本文提出Wahkon,一种将柯尔莫戈洛夫叠加定理与RKHS正则化结合的深层超叠加网络。该方法基于瓦尔巴的平滑样条传统,导出有限维深度表示定理,使训练可行并实现显式的逐层复杂度控制。我们证明了惩罚估计器正是在分层高斯过程先验下的最大后验估计,将样条与高斯过程对偶性扩展至深层复合结构。通过度量熵分析,我们在温和光滑条件下建立了最小最大最优收敛速率,并阐明深度与宽度如何与正则性权衡。实验表明,Wahkon在模拟基准和单细胞CITE-seq研究中均优于多层感知机、神经正切核及柯尔莫戈洛夫-阿诺德网络。通过统一柯尔莫戈洛夫叠加定理与RKHS正则化,Wahkon在一个框架内实现了准确性、可解释性与统计严谨性的统一。
原文摘要 · Abstract (English)
Deep learning excels at prediction but often lacks finite-sample guarantees and calibrated uncertainty; RKHS (Reproducing Kernel Hilbert Space)-based methods provide those guarantees but struggle to adapt in high dimensions. We propose Wahkon, a deep RKHS superposition network that unifies Kolmogorov's superposition principle with RKHS regularization in the smoothing-spline tradition of Wahba. This yields a finite-dimensional deep representer theorem that makes training tractable and provides explicit layerwise complexity control. We show the penalized estimator is exactly the MAP (maximum a posteriori) estimate under a hierarchical Gaussian-process prior, extending the spline/GP duality to deep compositions. Using metric-entropy arguments, we establish minimax-optimal convergence rates under mild smoothness and clarify how depth and width trade off with regularity. Empirically, Wahkon outperforms multilayer perceptrons, Neural Tangent Kernels, and Kolmogorov--Arnold Networks across simulation benchmarks and a single-cell CITE-seq study. By unifying Kolmogorov's superposition principle with RKHS regularization, Wahkon delivers accuracy, interpretability, and statistical rigor in a single framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。