用确定性方法分析模型泛化,分离几何与概率因素。
Separating Geometry from Probability in the Analysis of Generalization
- 通过优化问题对数据扰动的敏感性分析泛化性能
- 泛化界由样本间距离决定,不依赖独立同分布假设
- 适合关注理论严谨性与非统计范式的研究者
机器学习的目标是找到在未见数据上预测误差最小的模型。传统泛化分析假设样本来自无限总体且独立同分布,但此类统计假设无法验证。本文提出一种新视角:将泛化视为优化解对数据扰动的敏感性问题。在此框架下,泛化界可通过纯粹确定性方式推导,表现为变分原理,连接了样本内与样本外评估,误差项量化了样本外数据与样本内数据的接近程度。后续可结合统计假设,事后分析该误差项在平均或高概率意义下的大小。
原文摘要 · Abstract (English)
The goal of machine learning is to find models that minimize prediction error on data that has not yet been seen. Its operational paradigm assumes access to a dataset $S$ and articulates a scheme for evaluating how well a given model performs on an arbitrary sample. The sample can be $S$ (in which case we speak of ``in-sample'' performance) or some entirely new $S'$ (in which case we speak of ``out-of-sample'' performance). Traditional analysis of generalization assumes that both in- and out-of-sample data are i.i.d.\ draws from an infinite population. However, these probabilistic assumptions cannot be verified even in principle. This paper presents an alternative view of generalization through the lens of sensitivity analysis of solutions of optimization problems to perturbations in the problem data. Under this framework, generalization bounds are obtained by purely deterministic means and take the form of variational principles that relate in-sample and out-of-sample evaluations through an error term that quantifies how close out-of-sample data are to in-sample data. Statistical assumptions can then be used \textit{ex post} to characterize the situations when this error term is small (either on average or with high probability).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。