arXiv:2509.19788stat.MLcs.LG2025-09被引 1

提出新方法避免凸回归边界过拟合,可直接从数据估计误差上限。

Convex Regression with a Penalty

  • 在平方误差上界约束下,对次梯度施加惩罚以抑制边界过拟合
  • 证明了估计器及其次梯度在全空间一致几乎必然收敛,给出收敛速率
  • 适用于排队系统等待时间等真实场景,能自动确定误差上限

从 $n$ 个含噪观测值中估计未知的凸回归函数 $f_0: Ωsub R^d ightarrow R$ 的常用方法是寻找最小化平方误差和的凸函数。然而该估计器在 $Ω$ 边界处易出现过拟合,限制了实际应用。本文提出一种新估计器:在上界 $s_n$ 控制的平方误差和条件下,最小化次梯度上的惩罚项。关键优势在于 $s_n$ 可直接由数据估计。我们证明了该估计器及其次梯度在 $n o ty$ 时于 $Ω$ 上一致几乎必然收敛,并推导出收敛速率。通过在单服务器队列中估计等待时间展示了该方法的有效性。

原文摘要 · Abstract (English)

A common way to estimate an unknown convex regression function $f_0: Ω\subset \mathbb{R}^d \rightarrow \mathbb{R}$ from a set of $n$ noisy observations is to fit a convex function that minimizes the sum of squared errors. However, this estimator is known for its tendency to overfit near the boundary of $Ω$, posing significant challenges in real-world applications. In this paper, we introduce a new estimator of $f_0$ that avoids this overfitting by minimizing a penalty on the subgradient while enforcing an upper bound $s_n$ on the sum of squared errors. The key advantage of this method is that $s_n$ can be directly estimated from the data. We establish the uniform almost sure consistency of the proposed estimator and its subgradient over $Ω$ as $n \rightarrow \infty$ and derive convergence rates. The effectiveness of our estimator is illustrated through its application to estimating waiting times in a single-server queue.

凸回归统计学习误差控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。