证明深度RNN可逼近平滑序列函数,给出回归预测误差的最优保证。
Approximation Bounds for Recurrent Neural Networks with Application to Regression
- 构建RNN逼近依赖历史与当前信息的光滑函数
- 在独立同分布和指数β混合数据下达到最小最大误差率
- 为RNN回归提供统计性能保障,适合理论研究者
我们研究深度ReLU循环神经网络(RNN)的逼近能力,并探讨使用RNN进行非参数最小二乘回归的收敛性质。推导了RNN对霍尔德光滑函数的逼近误差上界,即RNN在每个时间步的输出可逼近仅依赖于过去和当前信息的函数(称为过去依赖函数)。通过精心设计,RNN可同时逼近一系列过去依赖的霍尔德函数。将这些逼近结果应用于回归问题,推导出经验风险最小化器预测误差的非渐近上界。在独立同分布和指数β-混合数据假设下,误差界均达到最小最大率,优于已有结果。结果为RNN的性能提供了统计保证。
原文摘要 · Abstract (English)
We study the approximation capacity of deep ReLU recurrent neural networks (RNNs) and explore the convergence properties of nonparametric least squares regression using RNNs. We derive upper bounds on the approximation error of RNNs for Hölder smooth functions, in the sense that the output at each time step of an RNN can approximate a Hölder function that depends only on past and current information, termed a past-dependent function. This allows a carefully constructed RNN to simultaneously approximate a sequence of past-dependent Hölder functions. We apply these approximation results to derive non-asymptotic upper bounds for the prediction error of the empirical risk minimizer in regression problem. Our error bounds achieve minimax optimal rate under both exponentially $β$-mixing and i.i.d. data assumptions, improving upon existing ones. Our results provide statistical guarantees on the performance of RNNs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。