arXiv:2605.08517cs.LGcs.CV2026-05被引 1

提出可评估混合学习与已知算子网络风险的深度风险估计算法。

A Deep Risk Estimator for Known Operator Learning

论文配图:A Deep Risk Estimator for Known Operator Learning
图 1 · 摘自论文原文
  • 基于训练误差上界,将总风险分解为各层贡献之和。
  • 替换学习层为已知算子可降低风险,样本需求与参数量成正比。
  • 适用于医学成像等物理信息神经网络,可预测所需训练数据量。

我们提出一种针对包含学习与已知算子混合结构的深度网络的统计风险估计算法。基于此前已知算子学习的极大训练误差上界,推导出连接网络期望误差与训练样本规模的深度风险估计器。该估计器将总风险分解为各学习层的贡献:已知算子贡献为零,学习层则包含受Barron工作启发的近似项和随样本数增加而减小的估计项。我们证明,当学习层被已知算子替代时,风险界会缩小,且对应样本需求与被替换层的可训练参数数量成正比。以计算机断层扫描为例,对比了算子感知滤波反投影网络与全连接替代模型(将整个重建流程压缩为单个学习稠密矩阵),预测的参数比例与解析分解中循环滤波器与稀疏反投影结构暴露的结构性稀疏完全一致。在小图像尺度下于CPU、中等图像尺度下于GPU均验证了相同缩放规律。该方法还可推广至硬编码已知物理操作的物理信息神经网络,对广泛开展算子感知深度学习的研究者具有参考价值。通过每轮迭代校准层间常数,估计界在所有训练集规模下与实测测试均方误差保持在两倍以内,因而可逆用于预测达到目标误差所需的训练样本数。

原文摘要 · Abstract (English)

We describe an approach for estimating the statistical risk of deep networks that contain a mix of learned and known operators. Building on the maximal training error bounds previously established for known operator learning, we derive a deep risk estimator that connects the expected error of a layered network to the size of the training sample. The estimator decomposes the total risk into a sum over learned layers; every known operator contributes zero to this sum, while every learned layer adds an approximation term inspired by Barron's classic work and an estimation term that decreases with the number of training samples. We are able to show that the bound shrinks whenever a learned layer is replaced by a known operator and that the corresponding sample requirement scales with the number of trainable parameters of the layer that is replaced. As an application, we use computed tomography as an example and compare an operator-aware filtered backprojection network with a fully connected substitute that collapses the entire reconstruction pipeline into a single learned dense matrix. The predicted parameter ratio coincides with the structural sparsity that the analytic decomposition into a circulant filter and a sparse backprojection exposes. We confirm the predicted scaling on CPU at small image scale and on GPU at medium image scale, all on the same scaling law. Beyond CT reconstruction, the estimator applies to physics-informed neural networks that hardcode a known physical operation in its architecture, and we expect the result to be of interest for a broad community working on operator-aware deep learning. Calibrating the per-layer constants on each sweep yields a bound that tracks the empirical test MSE within a factor of two at every training-set size, so the estimator can be inverted to predict how many training samples are required to reach a target error.

风险估计算子学习深度网络物理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。