arXiv:2507.06765cs.LG2025-07被引 1

提出新型平滑激活函数与扩散损失,提升非线性回归模型鲁棒性。

Robust Deep Network Learning of Nonlinear Regression Tasks by Parametric Leaky Exponential Linear Units (LELUs) and a Diffusion Metric

  • 设计可训练的平滑漏失指数线性单元(LELU),避免梯度消失
  • 新扩散损失有效检测过拟合,提升模型泛化能力
  • 适合需要高稳定性与抗敏感性的非线性回归任务

本文提出一种参数化激活函数,旨在改进多维非线性数据回归性能。众所周知,学习非线性数据需使用非线性激活函数。本工作表明,激活函数的光滑性与梯度特性对大型神经网络的过拟合与参数敏感性有显著影响。光滑但梯度消失的激活函数(如ELU、SiLU)性能受限,非光滑函数(如ReLU、Leaky-ReLU)则引入训练模型的不连续性。本文提出一种平滑的‘漏失指数线性单元’(LELU),具备非零梯度且可训练,实现性能提升。同时提出一种新型扩散损失度量,用于评估模型在过拟合方面的表现。

原文摘要 · Abstract (English)

This document proposes a parametric activation function (ac.f.) aimed at improving multidimensional nonlinear data regression. It is a established knowledge that nonlinear ac.f's are required for learning nonlinear datasets. This work shows that smoothness and gradient properties of the ac.f. further impact the performance of large neural networks in terms of overfitting and sensitivity to model parameters. Smooth but vanishing-gradient ac.f's such as ELU or SiLU (Swish) have limited performance and non-smooth ac.f's such as RELU and Leaky-RELU further impart discontinuity in the trained model. Improved performance is demonstrated with a smooth "Leaky Exponential Linear Unit", with non-zero gradient that can be trained. A novel diffusion-loss metric is also proposed to gauge the performance of the trained models in terms of overfitting.

非线性回归激活函数过拟合控制神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。