高维分位数预测新方法,理论强且自适应稀疏性。
A sparse PAC-Bayesian approach for high-dimensional quantile prediction
- 用缩放学生t先验与朗之万蒙特卡洛实现高效贝叶斯推断
- 理论证明预测误差达极小最优,且不依赖已知稀疏度
- 适合高维数据建模,尤其在变量远多于样本时表现佳
分位数回归是估计条件分位数的稳健方法,在计量经济学、统计学和机器学习中已取得显著进展。在高维情形下(协变量数量超过样本量),通过Lasso等正则化方法应对稀疏性挑战。虽然贝叶斯方法最初通过非对称拉普拉斯似然与分位数回归关联,但后验方差问题催生了伪似然/得分似然等新方法。本文提出一种新型概率机器学习框架用于高维分位数预测,采用缩放学生t先验与朗之万蒙特卡洛进行高效计算。该方法通过PAC-Bayes界建立了强理论保证,给出了非渐近的Oracle不等式,证明了预测误差达到极小最优并能自适应未知稀疏性。在模拟与真实数据上均表现出色,性能优于现有经典频率派与贝叶斯方法。
原文摘要 · Abstract (English)
Quantile regression, a robust method for estimating conditional quantiles, has advanced significantly in fields such as econometrics, statistics, and machine learning. In high-dimensional settings, where the number of covariates exceeds sample size, penalized methods like lasso have been developed to address sparsity challenges. Bayesian methods, initially connected to quantile regression via the asymmetric Laplace likelihood, have also evolved, though issues with posterior variance have led to new approaches, including pseudo/score likelihoods. This paper presents a novel probabilistic machine learning approach for high-dimensional quantile prediction. It uses a pseudo-Bayesian framework with a scaled Student-t prior and Langevin Monte Carlo for efficient computation. The method demonstrates strong theoretical guarantees, through PAC-Bayes bounds, that establish non-asymptotic oracle inequalities, showing minimax-optimal prediction error and adaptability to unknown sparsity. Its effectiveness is validated through simulations and real-world data, where it performs competitively against established frequentist and Bayesian techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。