用模拟推断解决选择偏差,无需可计算似然性。
Overcoming Selection Bias in Statistical Studies With Amortized Bayesian Inference

- 将选择机制嵌入生成器,实现无似然性下的贝叶斯推断
- 在三种不同选择机制下均获得校准良好的后验分布
- 适合复杂模型中存在观测缺失依赖未观测变量的场景
选择偏差发生在观测进入数据集的概率与感兴趣变量相关时,导致估计和不确定性量化出现系统性偏差。例如,在流行病学或调查研究中,某些结果个体更可能被纳入,从而造成偏倚的患病率估计,并产生显著下游影响。传统修正方法(如逆概率加权或显式似然模型)依赖于可计算的似然函数,限制了其在具有潜在动态或高维结构的复杂随机模型中的应用。基于模拟的推断可在无显式似然时进行贝叶斯分析,但通常假设缺失随机,因此当选择依赖于未观测结果或协变量时失效。本文提出一种包含偏差意识的模拟推断框架,将选择机制直接嵌入生成模拟器中,实现无需可计算似然的摊销贝叶斯推断。通过将选择偏差重构为模拟过程的一部分,该方法既能获得去偏估计,又能明确检测偏差存在性。框架集成诊断工具以检测模拟数据与观测数据之间的差异,并评估后验校准情况。在三个具有多样化选择机制的统计应用中,该方法均恢复了良好校准的后验分布,而基于似然的方法则产生偏倚估计。这些结果将选择偏差校正重铸为一个模拟问题,确立基于模拟的推断是应对选择偏差时实用且可验证的参数估计策略。
原文摘要 · Abstract (English)
Selection bias arises when the probability that an observation enters a dataset depends on variables related to the quantities of interest, leading to systematic distortions in estimation and uncertainty quantification. For example, in epidemiological or survey settings, individuals with certain outcomes may be more likely to be included, resulting in biased prevalence estimates with potentially substantial downstream impact. Classical corrections, such as inverse-probability weighting or explicit likelihood-based models of the selection process, rely on tractable likelihoods, which limits their applicability in complex stochastic models with latent dynamics or high-dimensional structure. Simulation-based inference enables Bayesian analysis without tractable likelihoods but typically assumes missingness at random and thus fails when selection depends on unobserved outcomes or covariates. Here, we develop a bias-aware simulation-based inference framework that explicitly incorporates selection into neural posterior estimation. By embedding the selection mechanism directly into the generative simulator, the approach enables amortized Bayesian inference without requiring tractable likelihoods. This recasting of selection bias as part of the simulation process allows us to both obtain debiased estimates and explicitly test for the presence of bias. The framework integrates diagnostics to detect discrepancies between simulated and observed data and to assess posterior calibration. The method recovers well-calibrated posterior distributions across three statistical applications with diverse selection mechanisms, including settings in which likelihood-based approaches yield biased estimates. These results recast the correction of selection bias as a simulation problem and establish simulation-based inference as a practical and testable strategy for parameter estimation under selection bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。