arXiv:2601.22326cs.LGstat.AP2026-01

用分层重要性采样高效监控低错误率分类模型,节省标签成本。

Label-Efficient Monitoring of Classification Models via Stratified Importance Sampling

  • 基于分层重要性采样,无需精确分布即可高效估计模型性能。
  • 在固定标签预算下,相比传统方法显著降低误差,提升监控效率。
  • 适合生产环境中的低错误率模型监控,尤其标签资源稀缺时使用。

在严格标签预算、一次性批量获取标签和极低错误率的条件下,监控分类模型的性能至关重要却极具挑战。本文提出一种基于分层重要性采样(SIS)的通用框架,直接应对这些约束。尽管SIS此前仅应用于特定领域,但我们的理论分析证明其广泛适用于分类模型监控。在温和条件下,SIS可生成无偏估计,并在有限样本下实现比重要性采样(IS)和分层随机采样(SRS)更优的均方误差(MSE)。该框架不依赖最优提议分布或分层划分:即使使用噪声代理和次优分层,仍能提升估计效率,但严重分布偏差会限制增益。跨二分类与多分类任务的实验表明,在固定标签预算下始终实现效率提升,验证了SIS作为原理严谨、标签高效且操作轻量的部署后模型监控方法的可行性。

原文摘要 · Abstract (English)

Monitoring the performance of classification models in production is critical yet challenging due to strict labeling budgets, one-shot batch acquisition of labels and extremely low error rates. We propose a general framework based on Stratified Importance Sampling (SIS) that directly addresses these constraints in model monitoring. While SIS has previously been applied in specialized domains, our theoretical analysis establishes its broad applicability to the monitoring of classification models. Under mild conditions, SIS yields unbiased estimators with strict finite-sample mean squared error (MSE) improvements over both importance sampling (IS) and stratified random sampling (SRS). The framework does not rely on optimally defined proposal distributions or strata: even with noisy proxies and sub-optimal stratification, SIS can improve estimator efficiency compared to IS or SRS individually, though extreme proposal mismatch may limit these gains. Experiments across binary and multiclass tasks demonstrate consistent efficiency improvements under fixed label budgets, underscoring SIS as a principled, label-efficient, and operationally lightweight methodology for post-deployment model monitoring.

模型监控重要性采样标签效率生产部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。