arXiv:2607.07637cs.LGcs.NA2026-07

基于误差估计动态调整网络深度,让模型更高效捕捉复杂问题。

An optimal control approach for neural network architecture adaptation with a posteriori error estimation

论文配图:An optimal control approach for neural network architecture adaptation with a posteriori error estimation
图 1 · 摘自论文原文
  • 将网络训练视为连续最优控制问题,用误差分布指导加层位置。
  • 在纳维-斯托克斯方程上实现更好泛化性能,优于现有方法。
  • 理论严格,适合需要高精度与可解释性的科学计算场景。

本文提出一种基于后验误差估计的神经网络深度自适应新方法。通过将神经网络训练建模为连续时间最优控制问题,推导出严格的误差估计,量化各层近似误差分布。该误差分解使网络可在误差最大处精准插入新层,有效捕捉底层问题的非线性变化。框架引入一种新型网络结构,将权重和偏置视为分段线性函数,误差估计算器约束离散表示与真实连续最优解之间的差异。方法借鉴有限元分析中的对偶加权残差法,给出功能误差的可计算上界。关键理论贡献在于推导出显式误差界,将总近似误差分解为区间贡献,为定向架构优化提供严格依据。在科学数据集上的实验表明,该方法在学习纳维-斯托克斯方程的可观测量到参数映射任务中,显著优于现有架构自适应方法,泛化性能持续领先。

原文摘要 · Abstract (English)

This work presents a novel approach for adapting neural network architecture along the depth based on a posteriori error estimation. By formulating neural network training as a continuous-time optimal control problem, we derive rigorous error estimates that quantify how approximation error distributes across network layers. This error decomposition enables a principled depth adaptation strategy: new layers are inserted at locations of maximum estimated error, allowing the network to efficiently capture complex, nonlinear variations in the underlying problem. Our framework introduces a novel network architecture that treats weights and biases as piecewise linear functions varying across layers, with the error estimator bounding the discrepancy between this discrete representation and the true continuous optimal control solution. The approach leverages dual weighted residual methodology from finite element analysis to derive computable upper bounds on the functional error. A key theoretical contribution is the derivation of explicit error bounds that decompose the total approximation error into interval-wise contributions, providing a rigorous basis for targeted architecture refinement. We demonstrate the effectiveness of our method on scientific datasets, including learning the observable-to-parameter map for the Navier-Stokes equation. Numerical results reveal that our approach consistently outperforms existing architecture adaptation methods in terms of generalization performance.

神经网络优化误差估计科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。