用扩散模型统一生成与回归,提升复杂数据的非参数建模精度。
Deep Bootstrap
- 基于条件扩散模型学习响应变量分布,生成新样本。
- 理论证明在Wasserstein距离下收敛速度最优。
- 适合高维、多模态数据的回归任务,可扩展性强。
本文提出一种基于条件扩散模型的新型深度自举框架,用于非参数回归。通过构建条件扩散模型学习给定协变量下响应变量的分布,并将原始协变量与新生成的响应配对以生成自举样本。我们将非参数回归重新表述为条件样本均值估计,直接通过学习到的条件扩散模型实现。与传统自举方法中分离估计、采样和回归的流程不同,本方法将三者整合进统一的生成框架。得益于扩散模型的强大表达能力,该方法可高效从高维或多重模态分布中采样,并实现精确的非参数估计。我们建立了严格的理论保证,推导出学习到的条件分布与目标分布间在Wasserstein距离下的最优端到端收敛速率。在此基础上,进一步证明了所提自举程序的收敛性。数值实验表明,该方法在复杂回归任务中具有高效性和可扩展性。
原文摘要 · Abstract (English)
In this work, we propose a novel deep bootstrap framework for nonparametric regression based on conditional diffusion models. Specifically, we construct a conditional diffusion model to learn the distribution of the response variable given the covariates. This model is then used to generate bootstrap samples by pairing the original covariates with newly synthesized responses. We reformulate nonparametric regression as conditional sample mean estimation, which is implemented directly via the learned conditional diffusion model. Unlike traditional bootstrap methods that decouple the estimation of the conditional distribution, sampling, and nonparametric regression, our approach integrates these components into a unified generative framework. With the expressive capacity of diffusion models, our method facilitates both efficient sampling from high-dimensional or multimodal distributions and accurate nonparametric estimation. We establish rigorous theoretical guarantees for the proposed method. In particular, we derive optimal end-to-end convergence rates in the Wasserstein distance between the learned and target conditional distributions. Building on this foundation, we further establish the convergence guarantees of the resulting bootstrap procedure. Numerical studies demonstrate the effectiveness and scalability of our approach for complex regression tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。