从逼近理论视角审视机器学习,揭示模型泛化与理论的脱节。
An Approximation Theory Perspective on Machine Learning
- 用逼近理论分析神经网络与核方法的表达能力
- 指出当前框架在泛化性能上的理论缺陷
- 提出无需显式学习流形特征的新逼近方法
机器学习的核心问题常被表述为:给定来自未知概率分布的样本集 $\\(\{(x_j, y_j)\ }_{j=1}^M$,目标是构建一个函数模型 $f$,使得对任意从相同分布中抽取的 $(x, y)$ 都有 $f(x) \approx y$。神经网络和基于核的方法因其高效的并行计算能力而被广泛采用。过去35年中,这些方法的逼近能力(即表达力)已得到深入研究。本文回顾该领域关键思想,探讨浅层/深层网络、流形上的逼近、物理信息神经代理、神经算子及Transformer架构等新兴趋势。尽管函数逼近是机器学习的基础问题,但逼近理论并未在该领域的理论基础中占据核心地位。这一脱节导致训练模型在未见或无标签数据上的泛化性能往往不明确。本文分析当前机器学习框架的局限性及其与逼近理论间差距的原因,并引入新研究:在无需学习特定流形特征(如拉普拉斯-贝尔特拉米算子的特征分解或图册构造)的前提下,实现对未知流形上的函数逼近。在许多机器学习任务中,尤其是分类任务,标签 $y_j$ 来自有限值集合。
原文摘要 · Abstract (English)
A central problem in machine learning is often formulated as follows: Given a dataset $\{(x_j, y_j)\}_{j=1}^M$, which is a sample drawn from an unknown probability distribution, the goal is to construct a functional model $f$ such that $f(x) \approx y$ for any $(x, y)$ drawn from the same distribution. Neural networks and kernel-based methods are commonly employed for this task due to their capacity for fast and parallel computation. The approximation capabilities, or expressive power, of these methods have been extensively studied over the past 35 years. In this paper, we will present examples of key ideas in this area found in the literature. We will discuss emerging trends in machine learning including the role of shallow/deep networks, approximation on manifolds, physics-informed neural surrogates, neural operators, and transformer architectures. Despite function approximation being a fundamental problem in machine learning, approximation theory does not play a central role in the theoretical foundations of the field. One unfortunate consequence of this disconnect is that it is often unclear how well trained models will generalize to unseen or unlabeled data. In this review, we examine some of the shortcomings of the current machine learning framework and explore the reasons for the gap between approximation theory and machine learning practice. We will then introduce our novel research to achieve function approximation on unknown manifolds without the need to learn specific manifold features, such as the eigen-decomposition of the Laplace-Beltrami operator or atlas construction. In many machine learning problems, particularly classification tasks, the labels $y_j$ are drawn from a finite set of values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。