提出一种鲁棒非参数回归方法,提升模型在分布不确定下的泛化能力。
Wasserstein Distributionally Robust Nonparametric Regression
- 基于Wasserstein距离阶数设计正则化机制,区分平滑与梯度约束。
- 理论证明估计器收敛率达 $n^{-2β/(d+2β)}$,对高维情形为最优。
- 适用于回归与分类,实测在MNIST上展现强鲁棒性。
Wasserstein分布鲁棒优化(WDRO)通过在给定模糊集内最小化局部最坏风险,增强模型不确定性下的统计学习性能。尽管WDRO在参数框架中被广泛研究,其在非参数设置中的理论性质仍不充分。本文研究非参数回归中的WDRO,首次揭示了水恩斯坦距离阶数 $k$ 的结构性差异:$k=1$ 导致Lipschitz型正则化,而 $k > 1$ 对应梯度范数正则化。为应对模型误设,分析了超出局部最坏风险的误差,推导出基于范数约束前馈神经网络估计器的非渐近误差界。该分析依赖于新的覆盖数与逼近界,同时控制函数及其梯度。所提估计器收敛率可达 $n^{-2β/(d+2β)}$(含对数因子),其中 $β$ 取决于目标函数光滑性与网络参数。此速率在高维常见条件下被证明为极小极大最优。此外,该最坏风险界可推出自然风险的保证,确保对模糊集内任意分布的鲁棒性。框架还展示在回归与分类问题中的通用性。模拟实验与在MNIST数据集上的应用进一步验证了估计器的鲁棒性。
原文摘要 · Abstract (English)
Wasserstein distributionally robust optimization (WDRO) strengthens statistical learning under model uncertainty by minimizing the local worst-case risk within a prescribed ambiguity set. Although WDRO has been extensively studied in parametric settings, its theoretical properties in nonparametric frameworks remain underexplored. This paper investigates WDRO for nonparametric regression. We first establish a structural distinction based on the order $k$ of the Wasserstein distance, showing that $k=1$ induces Lipschitz-type regularization, whereas $k > 1$ corresponds to gradient-norm regularization. To address model misspecification, we analyze the excess local worst-case risk, deriving non-asymptotic error bounds for estimators constructed using norm-constrained feedforward neural networks. This analysis is supported by new covering number and approximation bounds that simultaneously control both the function and its gradient. The proposed estimator achieves a convergence rate of $n^{-2β/(d+2β)}$ up to logarithmic factors, where $β$ depends on the target's smoothness and network parameters. This rate is shown to be minimax optimal under conditions commonly satisfied in high-dimensional settings. Moreover, these bounds on the excess local worst-case risk imply guarantees on the excess natural risk, ensuring robustness against any distribution within the ambiguity set. We show the framework's generality across regression and classification problems. Simulation studies and an application to the MNIST dataset further illustrate the estimator's robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。