揭示高维学习中模型崩溃的临界样本量,提出可验证的稳定性判据。
Spectral Thresholds for Identifiability and Stability:Finite-Sample Phase Transitions in High-Dimensional Learning
- 基于费雪信息矩阵最小特征值,建立有限样本稳定性阈值
- 当样本量低于 $d/n$ 临界比例时,模型稳定性和可识别性会突然失效
- 适用于高维统计推断与模型诊断,尤其适合小样本场景
在高维学习中,模型在样本量降至某一临界水平前保持稳定,一旦低于该水平则会突然崩溃。这种不稳定性并非算法特有,而是源于几何机制:当最弱的费雪特征方向低于样本波动水平时,参数可识别性即告失效。本文提出的费雪阈值定理证明,稳定性要求最小费雪特征值超过一个明确的 $O(\ ext{\sqrt{d/n}})$ 量级的界限。该阈值为有限样本下必要且普适的条件,标志着可靠集中与必然失败之间的尖锐相变。为使原理可操作,我们引入费雪底限(Fisher floor),一种对平滑和预处理鲁棒的谱正则化方法。高斯混合与逻辑回归的合成实验验证了预测的相变行为,符合 $d/n$ 标度规律。统计上,该阈值将经典特征值条件提升为非渐近律;学习理论上,它定义了谱样本复杂度前沿,实现了理论与诊断工具的融合。
原文摘要 · Abstract (English)
In high-dimensional learning, models remain stable until they collapse abruptly once the sample size falls below a critical level. This instability is not algorithm-specific but a geometric mechanism: when the weakest Fisher eigendirection falls beneath sample-level fluctuations, identifiability fails. Our Fisher Threshold Theorem formalizes this by proving that stability requires the minimal Fisher eigenvalue to exceed an explicit $O(\sqrt{d/n})$ bound. Unlike prior asymptotic or model-specific criteria, this threshold is finite-sample and necessary, marking a sharp phase transition between reliable concentration and inevitable failure. To make the principle constructive, we introduce the Fisher floor, a verifiable spectral regularization robust to smoothing and preconditioning. Synthetic experiments on Gaussian mixtures and logistic models confirm the predicted transition, consistent with $d/n$ scaling. Statistically, the threshold sharpens classical eigenvalue conditions into a non-asymptotic law; learning-theoretically, it defines a spectral sample-complexity frontier, bridging theory with diagnostics for robust high-dimensional inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。