arXiv:2608.22597stat.MLcs.LG2026-08

提出一种不随数据尺度变化的稀疏模型采样方法,减少罕见事件数据的信息损失。

Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models

  • 基于自适应Lasso设计预测误差最小化的采样函数
  • 在模拟和真实数据上验证了采样后预测精度显著提升
  • 适合处理含大量无效特征的罕见事件数据

子采样在应对大规模罕见事件数据的计算挑战时非常有效。过度激进的采样会损害估计效率,因此最优采样至关重要以减少信息损失。然而,现有最优采样概率依赖于数据尺度,不当的缩放可能造成低效采样。这一问题在存在无效特征时尤为严重,因为它们对采样概率的影响可能被任意放大。本文针对稀疏模型场景,提出一种尺度不变的最优采样函数,不关注参数估计本身,而是以最小化预测误差为目标。通过自适应Lasso构建估计流程并建立其oracle性质,验证了采样的合理性。进一步推导出最小化加权逆概率(IPW)自适应Lasso预测误差的尺度不变最优采样函数,并提出基于最大采样条件似然(MSCL)的估计器以提升估计效率。在模拟与真实数据集上的实验表明,所提方法性能优越。

原文摘要 · Abstract (English)

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.

稀疏模型罕见事件采样优化自适应Lasso

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。