arXiv:2410.05071cs.LGcs.SY2024-10被引 1

随机浅层ReLU网络可同时高精度逼近函数及其梯度,适用于控制中的策略评估。

Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications

  • 用随机初始化的浅层ReLU网络逼近函数与梯度。
  • 梯度误差为O((log m/m)^1/2),优于此前结果。
  • 适合需函数与梯度联合逼近的连续时间控制问题。

神经网络广泛用于控制中未知函数的逼近。常见架构为单隐层(即浅层网络),其中输入参数预先固定,仅训练输出参数。传统理论分析指出:若存在输出参数可足够精确逼近未知函数,则可实现期望控制性能。但长期存在的理论缺口在于:无法保证对固定输入参数,通过训练输出参数能达到所需精度。我们近期工作部分填补了这一空白,证明若输入参数随机选取,则对任意充分光滑函数,以高概率存在输出参数使逼近误差为O((1/m)^1/2),其中m为神经元数量。然而,某些应用如连续时间价值函数逼近,要求网络同时以足够精度逼近未知函数及其梯度。本文表明,当输入参数随机生成且输出参数可训练时,梯度误差为O((log m/m)^1/2),并改进了先前工作的常数项。我们进一步展示了该结果在策略评估问题中的应用。

原文摘要 · Abstract (English)

Neural networks are widely used to approximate unknown functions in control. A common neural network architecture uses a single hidden layer (i.e. a shallow network), in which the input parameters are fixed in advance and only the output parameters are trained. The typical formal analysis asserts that if output parameters exist to approximate the unknown function with sufficient accuracy, then desired control performance can be achieved. A long-standing theoretical gap was that no conditions existed to guarantee that, for the fixed input parameters, required accuracy could be obtained by training the output parameters. Our recent work has partially closed this gap by demonstrating that if input parameters are chosen randomly, then for any sufficiently smooth function, with high-probability there are output parameters resulting in $O((1/m)^{1/2})$ approximation errors, where $m$ is the number of neurons. However, some applications, notably continuous-time value function approximation, require that the network approximates the both the unknown function and its gradient with sufficient accuracy. In this paper, we show that randomly generated input parameters and trained output parameters result in gradient errors of $O((\log(m)/m)^{1/2})$, and additionally, improve the constants from our prior work. We show how to apply the result to policy evaluation problems.

神经网络梯度逼近控制应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。