一种无需假设的核密度学习框架,可准确估计概率测度间的密度变化。
Kernel Density Machines
- 基于核方法构建,不依赖传统非参数估计的结构假设。
- 证明了样本估计的一致性与泛函中心极限定理,误差率最优。
- 适用于两样本检验与条件分布估计,尤其在高维场景表现优异。
我们提出核密度机器(KDM),一种无偏见的基于核的方法,用于在最小假设下学习概率测度之间的Radon-Nikodym导数(即密度)。该方法适用于一般可测空间,避免了经典非参数密度估计器常见的结构约束。我们构造了一个样本估计量,并证明其一致性及泛函中心极限定理。为提升可扩展性,开发了类似Nyström的低秩近似方法,并推导出最优误差率,填补了密度学习领域缺乏此类保证的空白。通过在核基两样本检验和条件分布估计中的应用,展示了KDM的通用性;后者在维度无关的保证下优于局部平滑方法。模拟与真实数据实验表明,KDM在准确性、可扩展性及多任务表现上均具竞争力。
原文摘要 · Abstract (English)
We introduce kernel density machines (KDM), an agnostic kernel-based framework for learning the Radon-Nikodym derivative (density) between probability measures under minimal assumptions. KDM applies to general measurable spaces and avoids the structural requirements common in classical nonparametric density estimators. We construct a sample estimator and prove its consistency and a functional central limit theorem. To enable scalability, we develop Nystrom-type low-rank approximations and derive optimal error rates, filling a gap in the literature where such guarantees for density learning have been missing. We demonstrate the versatility of KDM through applications to kernel-based two-sample testing and conditional distribution estimation, the latter enjoying dimension-free guarantees beyond those of locally smoothed methods. Experiments on simulated and real data show that KDM is accurate, scalable, and competitive across a range of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。