提出新不等式,解决无界目标函数的随机优化误差问题。
Concentration Inequalities for Stochastic Optimization of Unbounded Objective Functions with Application to Denoising Score Matching
- 基于样本依赖的均值差界改进麦迪亚姆德不等式。
- 首次为无界分布提供统一的大数定律与泛化误差界。
- 适用于去噪得分匹配等需重用样本的算法,揭示其优势。
我们推导出一类新型浓度不等式,用于界定一大类随机优化问题的统计误差,重点处理无界目标函数情形。核心工具包括:1)一种基于样本依赖单分量均值差界的新型麦迪亚姆德不等式,由此得出无界函数的新型一致大数定律;2)一类满足样本依赖Lipschitz性质的函数族的新Rademacher复杂度界,可推广至具有无界支撑的广泛分布。作为应用,我们为去噪得分匹配(Denoising Score Matching, DSM)推导了统计误差界,该方法天然涉及无界目标函数和无界支撑分布,即使数据分布本身有界亦然。此外,我们的结果量化了在使用易采样的辅助随机变量(如DSM中的高斯变量)时样本重用的收益。
原文摘要 · Abstract (English)
We derive novel concentration inequalities that bound the statistical error for a large class of stochastic optimization problems, focusing on the case of unbounded objective functions. Our derivations utilize the following key tools: 1) A new form of McDiarmid's inequality that is based on sample-dependent one-component mean-difference bounds and which leads to a novel uniform law of large numbers result for unbounded functions. 2) A new Rademacher complexity bound for families of functions that satisfy an appropriate sample-dependent Lipschitz property, which allows for application to a large class of distributions with unbounded support. As an application of these results, we derive statistical error bounds for denoising score matching (DSM), an application that inherently requires one to consider unbounded objective functions and distributions with unbounded support, even in cases where the data distribution has bounded support. In addition, our results quantify the benefit of sample-reuse in algorithms that employ easily-sampled auxiliary random variables in addition to the training data, e.g., as in DSM, which uses auxiliary Gaussian random variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。