arXiv:2411.02784stat.MLcs.LG2024-11被引 4

提出更紧的RNN泛化误差界,提升模型可信度。

Generalization and Risk Bounds for Recurrent Neural Networks

  • 基于统一框架计算Rademacher复杂度,适配多种损失函数。
  • 使用ramp损失时,新界比现有方法更紧,平均提升13.80%(tanh)和3.01%(ReLU)。
  • 适用于多分类任务,对经验风险最小化估计器给出精确误差界。

循环神经网络(RNNs)在序列数据预测中取得显著成功,但其理论研究仍滞后于实际应用,主要因其复杂的互联结构。本文为标准RNNs建立了新的泛化误差界,并提出一个统一的Rademacher复杂度计算框架,可适用于多种损失函数。当采用ramp损失时,我们的界在相同关于权值矩阵Frobenius范数与谱范数的假设及若干温和条件下,优于现有方法。数值实验表明,在三个公开数据集上,本方法得到的泛化界是现有最紧的。相比次紧的界,平均改进幅度分别为13.80%(tanh激活)和3.01%(ReLU激活)。此外,当损失函数满足Bernstein条件时,我们还推导出基于ERM的RNN估计器在多分类问题中的精确估计误差界。

原文摘要 · Abstract (English)

Recurrent Neural Networks (RNNs) have achieved great success in the prediction of sequential data. However, their theoretical studies are still lagging behind because of their complex interconnected structures. In this paper, we establish a new generalization error bound for vanilla RNNs, and provide a unified framework to calculate the Rademacher complexity that can be applied to a variety of loss functions. When the ramp loss is used, we show that our bound is tighter than the existing bounds based on the same assumptions on the Frobenius and spectral norms of the weight matrices and a few mild conditions. Our numerical results show that our new generalization bound is the tightest among all existing bounds in three public datasets. Our bound improves the second tightest one by an average percentage of 13.80% and 3.01% when the $\tanh$ and ReLU activation functions are used, respectively. Moreover, we derive a sharp estimation error bound for RNN-based estimators obtained through empirical risk minimization (ERM) in multi-class classification problems when the loss function satisfies a Bernstein condition.

RNN泛化误差机器学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。