提出AOPU单元,提升软传感器训练稳定性与可解释性。
Approximated Orthogonal Projection Unit: Stabilizing Regression Network Training Using Natural Gradient
- 通过双参数截断梯度反向传播,逼近自然梯度
- 理论证明可实现最小方差估计,收敛更稳定
- 适合工业软传感器场景,尤其关注在线优化的领域
神经网络在先进软传感器模型中广泛应用,因其具备特征提取与函数逼近能力。当前研究多聚焦于模型离线精度,但在工业软传感器场景中,线上优化稳定性与可解释性优先于精度。为此,本文提出一种新神经网络结构——近似正交投影单元(AOPU),具有坚实的数学基础,显著提升训练稳定性。AOPU在双参数处截断梯度反向传播,优化可追踪参数更新,增强训练鲁棒性。进一步证明,AOPU在神经网络中可实现最小方差估计(MVE),其截断梯度近似自然梯度(NG)。在两个化工过程数据集上的实证结果表明,AOPU在实现稳定收敛方面优于其他模型,为软传感器领域带来重要进展。
原文摘要 · Abstract (English)
Neural networks (NN) are extensively studied in cutting-edge soft sensor models due to their feature extraction and function approximation capabilities. Current research into network-based methods primarily focuses on models' offline accuracy. Notably, in industrial soft sensor context, online optimizing stability and interpretability are prioritized, followed by accuracy. This requires a clearer understanding of network's training process. To bridge this gap, we propose a novel NN named the Approximated Orthogonal Projection Unit (AOPU) which has solid mathematical basis and presents superior training stability. AOPU truncates the gradient backpropagation at dual parameters, optimizes the trackable parameters updates, and enhances the robustness of training. We further prove that AOPU attains minimum variance estimation (MVE) in NN, wherein the truncated gradient approximates the natural gradient (NG). Empirical results on two chemical process datasets clearly show that AOPU outperforms other models in achieving stable convergence, marking a significant advancement in soft sensor field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。