arXiv:2606.23295cs.LGmath.PR2026-06

为统计学习的最小风险提供非渐近高置信度上下界

Non-asymptotic estimates of the minimal risk in statistical learning

  • 基于泛函积分条件放宽经典有界性假设
  • 最小风险上下界在小样本下仍保持高置信度
  • 适用于检测学习模型缺陷,适合理论研究者

本文证明了统计学习中经验风险原则(ERP)两类误差概率的集中不等式,给出了最小风险(以最小经验风险表示)的非渐近高置信度上下界。传统经验风险函数有界性假设被放松为高斯或指数可积性条件。最小风险下界置信度与训练参数数量及输入向量维度无关,可高效检测学习机缺陷;上界置信度在样本量n远大于参数集Θ在Orlicz度量d_{ψ_1}下的盒维数时可保证较高。研究基于Talagrand集中不等式(Bousquet与Klein-Rio的精化版本)、运输熵不等式以及经验过程与统计学习理论的最新进展。

原文摘要 · Abstract (English)

In this paper we prove some concentration inequalities for two types of error probabilities in the Empirical Risk Principle (ERP) in statistical learning, which provide a lower bound and an upper bound for the minimal risk (in terms of the minimal empirical risk) with non-asymptotic high confidence. The usual boundedness condition of the empirical risk function is relaxed to the Gaussian or exponential integrability condition. The confidence of the lower bound of the minimal risk is shown to be independent of the number of training parameters and the dimension of the input vectors, allowing one to detect the deficiency of a learning machine efficiently; and the confidence of the upper bound of the minimal risk is proved to be high provided that the sample size $n$ is much greater than the box dimension of the parameter set $Θ$ in the Orlicz metric $d_{ψ_1}$ associated with the risk functions. Our work is based on Talagrand's concentration inequalities (the sharp versions by Bousquet and Klein-Rio), transport-entropy inequalities and the recent progress in the theory of empirical processes and statistical learning.

统计学习风险估计集中不等式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。