arXiv:2605.13127stat.MLcs.LG2026-05被引 1

用小波构造新DPP核,提升小批量采样效率与精度。

State-of-art minibatches via novel DPP kernels: discretization, wavelets, and rough objectives

  • 基于小波构建连续域DPP核,理论性能优于现有方法。
  • 提出离散化方法,保持方差衰减且支持高效采样。
  • 适用于低光滑度目标函数,适合追求高鲁棒性的模型训练。

确定性点过程(DPP)作为独立采样的核化替代方案,在生成高效小批量、核心集等大规模数据的紧凑表示方面展现出潜力。尽管已有理论基础和良好实证表现,当前基于DPP的核心集或小批量仍面临两大挑战:一是需要具备特定方差缩减性质的DPP族,通常在连续空间中构造,但已知实例稀少;二是需对给定数据集构造一个离散DPP,继承此类方差缩减特性。本文在推动DPP成为机器学习子采样工具箱的进程中,从两方面推进:首先,提出基于小波的欧氏空间新DPP,其准确率保证优于现有最优结果;其次,引入通用方法,将更利于理论分析的连续DPP转化为离散核,同时保持期望方差衰减,并揭示离散核的低秩结构,使DPP采样计算成本显著降低。此外,扩展了可受益于DPP小批量与核心集的机器学习任务范围,涵盖任意低正则性目标函数,并提供显式适应该正则性的速率保证。

原文摘要 · Abstract (English)

Determinantal point processes (DPPs) have emerged as a kernelized alternative to vanilla independent sampling for generating efficient minibatches, coresets and other parsimonious representations of large-scale datasets. While theoretical foundations and promising empirical performance have been demonstrated, there are two challenges for current proposals for DPP-based coresets or minibatches. The first is the need for families of DPPs with certain key variance reduction properties, usually constructed in a continuous setting, of which there are few known examples. The second is the need for an ad-hoc construction of a discrete DPP defined on a given dataset, that inherits such variance reduction. In this work, we contribute to the programme of establishing DPPs as a subsampling toolbox for ML by advancing on these two fronts. First, we propose new DPPs on the Euclidean space based on wavelets, with provably better accuracy guarantees than the best known rates. Second, we introduce a general method to convert such continuous DPPs, which are more amenable to proving analytical statements, into discrete kernels, which are pertinent for subsampling tasks such as minibatch and coreset constructions. This conversion mechanism simultaneously preserves the desired variance decay and reveals a low-rank decomposition of the discrete kernel, which makes sampling the corresponding DPP computationally inexpensive. En route, we enlarge the class of ML tasks amenable to improvements via DPP-based minibatches and coresets to include objective functions with arbitrarily low regularity, and rate guarantees that explicitly adapt to this regularity.

DPP小批量采样小波核心集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。