arXiv:2411.00471stat.MEcs.LG2024-11

新方法提升线性模型变量选择效率,自动区分参数重要性。

Dirichlet process mixtures of block $g$ priors for model selection and prediction in linear models

  • 用分块g先验的狄利克雷过程混合建模,按数据自选参数块差异化收缩。
  • 小样本中大效应下,检测微弱显著效应的效能更高,假阳性略增。
  • 适合高维线性模型变量选择,尤其含少数强效应的场景。

本文提出狄利克雷过程混合分块g先验,用于线性模型的变量选择与预测。该方法扩展了传统g先验,允许对数据选定的参数块进行差异收缩,同时充分考虑预测变量的相关结构,连接了变量选择与连续收缩先验的研究脉络。我们证明该先验在多种意义下具有一致性,尤其避免了Som等(2016)指出的条件林德利悖论。此外,我们设计了一种马尔可夫链蒙特卡洛算法,仅需少量调参即可实现后验推断。在真实与模拟数据上的实证研究显示:当存在少量极大效应时,该先验能显著提升对较小但显著效应的检出能力,且假阳性数量增加极小。

原文摘要 · Abstract (English)

This paper introduces Dirichlet process mixtures of block $g$ priors for model selection and prediction in linear models. These priors are extensions of traditional mixtures of $g$ priors that allow for differential shrinkage for various (data-selected) blocks of parameters while fully accounting for the predictors' correlation structure, providing a bridge between the literatures on model selection and continuous shrinkage priors. We show that Dirichlet process mixtures of block $g$ priors are consistent in various senses and, in particular, that they avoid the conditional Lindley ``paradox'' highlighted by Som et al. (2016). Further, we develop a Markov chain Monte Carlo algorithm for posterior inference that requires only minimal ad-hoc tuning. Finally, we investigate the empirical performance of the prior in various real and simulated datasets. In the presence of a small number of very large effects, Dirichlet process mixtures of block $g$ priors lead to higher power for detecting smaller but significant effects without only a minimal increase in the number of false discoveries.

变量选择贝叶斯统计先验设计线性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。