新型神经算子可高效求解无穷多凸优化问题
Generative Neural Operators of Log-Complexity Can Simultaneously Solve Infinitely Many Convex Programs
- 用对数复杂度的生成平衡算子统一求解凸优化问题
- 参数量随精度提升仅对数增长,理论与实验一致
- 适合需快速求解大量相关优化问题的研究者
神经算子(NOs)是一类可同时求解无穷多个相关问题的深度学习模型,通过将问题映射到无限维空间进行操作。理论上,通用逼近定理表明此类模型需极大量参数,但实验却显示其表现良好。本文针对一类特定的生成平衡算子(GEOs),结合有限维深层均衡层,在可分希尔伯特空间上求解凸优化问题族。输入为光滑凸损失函数,输出为对应近似解。当输入损失函数属于合适的无限维紧集时,该GEO可以任意精度统一逼近解,且其秩、深度和宽度仅随逼近误差倒数的对数增长。我们进一步在三类应用中验证了理论结果与GEO的可训练性:(1) 非线性偏微分方程,(2) 随机最优控制问题,(3) 有流动性约束的金融对冲问题。
原文摘要 · Abstract (English)
Neural operators (NOs) are a class of deep learning models designed to simultaneously solve infinitely many related problems by casting them into an infinite-dimensional space, whereon these NOs operate. A significant gap remains between theory and practice: worst-case parameter bounds from universal approximation theorems suggest that NOs may require an unrealistically large number of parameters to solve most operator learning problems, which stands in direct opposition to a slew of experimental evidence. This paper closes that gap for a specific class of {NOs}, generative {equilibrium operators} (GEOs), using (realistic) finite-dimensional deep equilibrium layers, when solving families of convex optimization problems over a separable Hilbert space $X$. Here, the inputs are smooth, convex loss functions on $X$, and outputs are the associated (approximate) solutions to the optimization problem defined by each input loss. We show that when the input losses lie in suitable infinite-dimensional compact sets, our GEO can uniformly approximate the corresponding solutions to arbitrary precision, with rank, depth, and width growing only logarithmically in the reciprocal of the approximation error. We then validate both our theoretical results and the trainability of GEOs on three applications: (1) nonlinear PDEs, (2) stochastic optimal control problems, and (3) hedging problems in mathematical finance under liquidity constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。