重尾分布下,传统样本量法则失效,需考虑阈值共占用机制。
Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails
- 提出非高斯重尾设计中受限特征值的全新分析框架
- 证明在固定维数下,样本量需达Θ(β⁻¹[d log(1/β)+log(1/δ)])才能保证恢复
- 揭示路径级失败现象,适合高维统计与稳健估计研究者
受限特征值(RE)界决定范数正则化估计器的稳定恢复性能。对于各向同性次高斯测量,基准样本量为 $1 + w(A)^2$,其中 $w(A)$ 是归一化下降锥的高斯宽度。COLT 2015 开放问题(Banerjee 等, 2015)询问:仅基于均匀小球条件,是否仍可适用相同规律?我们以明确且系统的方式给出了否定回答:原命题在全维无依赖、任意集合形式下不成立,根本障碍在于同时阈值占用。一个恒定宽度的多面体下降锥,在固定小球常数下,其所有样本路径上的经验受限特征值均为零,直至环境维数的一半。更一般地,每个有限范围空间均可在任意窄球帽内实现精确阈值编码,并提升至完整多面体下降锥截面。对每个固定的阈值 VC 维 $d$,当 $eta o0$ 时,最坏情况下的样本复杂度为 $ heta(eta^{-1}[d\ abla(1/\beta) + \log(1/\delta)])$。该分离现象在精确各向同性和所有有限矩下依然存在:在同一恒定宽度锥上,高斯测量仅需 $O(1 + \log(1/\delta))$ 样本即成功,而各向同性重尾设计在 $n \lesssim \sqrt{p / \log p}$ 时路径级失败。高斯平滑产生处处正的 $C^\infty$ 密度,但仍保持极差的 RE。在各向同性下,一种分布无关的退路由仿射维数乘以平方包围半径决定,且对该类模型是紧的。
原文摘要 · Abstract (English)
Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gaussian measurements, the benchmark sample size is $1+w(A)^2$, where $w(A)$ is the Gaussian width of the normalized descent cone. The COLT 2015 open-problem note (Banerjee et al., 2015) asked whether the same law follows for heavy-tailed designs from a uniform small-ball condition alone. We give an explicit and systematic negative answer to the general question as formulated there: the proposed law fails in its full dimension-free, arbitrary-set form, and the missing obstruction is simultaneous threshold occupancy. A constant-width polyhedral descent cone with fixed small-ball constants has zero empirical RE on every sample path up to half the ambient dimension. More generally, every finite range space admits exact threshold encoding in an arbitrarily narrow spherical cap and a lift to a full polyhedral descent-cone section. For every fixed threshold VC dimension $d$, as $\beta\downarrow0$, the sharp worst-case sample complexity is $\Theta(\beta^{-1}[d\log(1/\beta)+\log(1/\delta)])$. The separation persists under exact isotropy and all finite moments: on the same constant-width cone, Gaussian measurements succeed with $O(1+\log(1/\delta))$ samples, whereas an isotropic heavy-tailed design fails pathwise for $n\lesssim\sqrt{p/\log p}$. Gaussian smoothing yields an everywhere-positive $C^\infty$ density while retaining arbitrarily poor RE. Under isotropy, a distribution-free fallback governed by affine dimension times squared enclosing radius is sharp on this family.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。