arXiv:2601.16041math.STcs.LG2026-01

约束越紧,估计风险反而可能越高,这在高噪声下尤其明显。

Risk reversal for least squares estimators under nested convex constraints

  • 通过嵌套凸集构造反例,揭示投影估计器的风险反转现象。
  • 在高噪声下,更紧的约束集导致更大的均方误差风险。
  • 该现象与噪声水平相关,适用于非高斯噪声和多种损失函数。

在受限随机优化中,通常认为只要可行集包含真实参数,缩小可行域不会增加估计器的统计风险。本文在高斯序列模型中证明这一直觉可能失效。给定紧凸集 $Θackslashsubseteq ackslashmathbb{R}^d$,观测数据 $Y = θ^ackslashstar + σZ$,其中 $Z \sim N(0, I_d)$,目标是估计未知参数 $θ^ackslashstar \in Θ$。此时最大似然估计与最小二乘估计(LSE)等价,即 $Y$ 到 $Θ$ 的欧氏投影。本文构造了明确反例,展示风险反转:当噪声足够大时,存在嵌套紧凸集 $Θ_S \subsetneq Θ_L$ 及 $θ^ackslashstar \in Θ_S$,使得 $Θ_S$ 上的 LSE 均方误差风险严格大于 $Θ_L$ 上的。进一步证明,该现象在最坏情形风险下仍成立。结果还表明,该现象不限于高斯噪声或平方损失。通过对比噪声尺度,发现当噪声趋于零时,风险由切锥的统计维数主导,无法出现风险反转;而当噪声发散时,风险依赖于约束集的整体几何结构,$Θ_S$ 在 $Θ_L$ 内的嵌入方式可导致风险排序逆转。这些结果揭示了最小二乘估计的一个此前未被认识的失效模式。

原文摘要 · Abstract (English)

In constrained stochastic optimization, one expects that restricting the feasible set, provided it still contains the true parameter, should not increase the statistical risk of the corresponding projection estimator. We show that this intuition can fail, even in basic settings. We investigate this phenomenon in the Gaussian sequence model. Given a compact, convex set $Θ\subseteq \mathbb{R}^d$, one observes \[ Y = θ^\star + σZ, \qquad Z \sim N(0, I_d), \] and seeks to estimate an unknown $θ^\star \in Θ$. Here, the maximum likelihood estimator over $Θ$ coincides with the least squares estimator (LSE), given by the Euclidean projection of $Y$ onto $Θ$. We construct an explicit example exhibiting \emph{risk reversal}: for sufficiently large noise, there exist nested compact convex sets $Θ_S \subsetneq Θ_L$ and $θ^\star \in Θ_S$ such that the LSE constrained to $Θ_S$ has strictly larger squared-error risk than the LSE constrained to $Θ_L$. Moreover, we demonstrate that risk reversal can persist at the level of worst-case risk. Finally, we show that the phenomenon is not specific to Gaussian noise or squared-error risk, extending our results beyond both settings. We clarify this phenomenon by contrasting noise regimes. In the vanishing-noise limit, the risk is governed at first order by the statistical dimension of the tangent cone, and risk reversal cannot occur at the leading $σ^2$ scale. In the diverging-noise regime, the risk instead depends on the global geometry of the constraint sets, and the embedding of $Θ_S$ within $Θ_L$ can reverse the risk ordering. These results reveal a previously unrecognized failure mode of the LSE. They demonstrate that in sufficiently noisy settings, tightening a constraint can paradoxically degrade statistical performance.

统计推断凸优化风险分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。