对比经典与生成式方法,评估条件分布估计的性能与效率。
Generative and Nonparametric Approaches for Conditional Distribution Estimation: Methods, Perspectives, and Comparative Evaluations
- 用降维与非参数平滑结合的方法估计条件分布
- 生成模型在 Wasserstein 距离上表现更优,计算成本更高
- 适合需要不确定性建模的统计建模与预测任务
条件分布推断是统计学中的基础问题,对预测、不确定性量化和概率建模至关重要。本文综述并比较了代表性方法,涵盖经典非参数方法与现代生成模型。包括 Hall 与 Yao (2005) 的单指标法,通过降维与一维累积条件分布函数的非参数平滑估计;FlexCode (Izbicki and Lee, 2017) 和 DeepCDE (Dalmasso et al., 2020) 的基展开方法,将条件密度估计转化为一系列非参数回归问题;以及基于深度生成架构的两种新方法:生成式条件分布采样器(Zhou et al., 2023)和条件去噪扩散概率模型(Fu et al., 2024;Yang et al., 2025)。采用统一评估框架进行系统比较,使用均方误差(条件均值与标准差)和 Wasserstein 距离作为性能指标。同时讨论各方法的灵活性与计算成本,揭示其各自优势与局限。
原文摘要 · Abstract (English)
The inference of conditional distributions is a fundamental problem in statistics, essential for prediction, uncertainty quantification, and probabilistic modeling. A wide range of methodologies have been developed for this task. This article reviews and compares several representative approaches spanning classical nonparametric methods and modern generative models. We begin with the single-index method of Hall and Yao (2005), which estimates the conditional distribution through a dimension-reducing index and nonparametric smoothing of the resulting one-dimensional cumulative conditional distribution function. We then examine the basis-expansion approaches, including FlexCode (Izbicki and Lee, 2017) and DeepCDE (Dalmasso et al., 2020), which convert conditional density estimation into a set of nonparametric regression problems. In addition, we discuss two recent generative simulation-based methods that leverage modern deep generative architectures: the generative conditional distribution sampler (Zhou et al., 2023) and the conditional denoising diffusion probabilistic model (Fu et al., 2024; Yang et al., 2025). A systematic numerical comparison of these approaches is provided using a unified evaluation framework that ensures fairness and reproducibility. The performance metrics used for the estimated conditional distribution include the mean-squared errors of conditional mean and standard deviation, as well as the Wasserstein distance. We also discuss their flexibility and computational costs, highlighting the distinct advantages and limitations of each approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。