用压缩数据集提升多变量分布建模的效率与精度。
Scalable Learning of Multivariate Distributions via Coresets
- 基于重要性采样构建多变量条件变换模型的压缩数据集
- 在高概率下保证似然值误差在(1±ε)内,保持模型准确性
- 适合处理复杂非线性关系的大规模数据,尤其适用于统计学习场景
高效且可扩展的非参数或半参数回归分析与密度估计对统计学和机器学习至关重要。然而现有方法难以应对大规模数据。本文提出一种新型多变量条件变换模型(MCTM)的压缩数据集构造方法,显著提升其可扩展性与训练效率。据我们所知,这是首个针对半参数分布模型的压缩数据集。通过重要性采样实现大幅数据压缩,在高概率下保证对数似然值处于(1±ε)的乘法误差范围内,从而维持模型统计精度。相比以往仅用于全参数模型的压缩技术,本方法在复杂分布与非线性关系存在但未完全理解的场景中展现出更强适应性。为解决对数项归一化带来的数值问题,采用基于输入数据凸包的几何近似方法,确保大规模数据下的稳定、准确推断。数值实验表明,该方法在处理大规模复杂数据集时计算效率显著提升,为统计与机器学习领域广泛应用奠定基础。
原文摘要 · Abstract (English)
Efficient and scalable non-parametric or semi-parametric regression analysis and density estimation are of crucial importance to the fields of statistics and machine learning. However, available methods are limited in their ability to handle large-scale data. We address this issue by developing a novel coreset construction for multivariate conditional transformation models (MCTMs) to enhance their scalability and training efficiency. To the best of our knowledge, these are the first coresets for semi-parametric distributional models. Our approach yields substantial data reduction via importance sampling. It ensures with high probability that the log-likelihood remains within multiplicative error bounds of $(1\pm\varepsilon)$ and thereby maintains statistical model accuracy. Compared to conventional full-parametric models, where coresets have been incorporated before, our semi-parametric approach exhibits enhanced adaptability, particularly in scenarios where complex distributions and non-linear relationships are present, but not fully understood. To address numerical problems associated with normalizing logarithmic terms, we follow a geometric approximation based on the convex hull of input data. This ensures feasible, stable, and accurate inference in scenarios involving large amounts of data. Numerical experiments demonstrate substantially improved computational efficiency when handling large and complex datasets, thus laying the foundation for a broad range of applications within the statistics and machine learning communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。