解决小数据下异常检测的精度与稳定性矛盾,提升真实场景检测能力
Between Resolution Collapse and Variance Inflation: Weighted Conformal Anomaly Detection in Low-Data Regimes
- 用连续加权核密度估计解耦局部适应与尾部分辨率
- 在小样本下恢复零发现问题的检测能力,统计功效显著提升
- 适合数据分布漂移严重、样本稀缺的工业异常检测场景
标准共形异常检测在可交换性假设下提供有限样本保证,但现实数据常存在分布漂移,需采用加权共形方法适应局部非平稳性。我们发现这种适应引发最小可达p值与其稳定性的关键权衡:重要性权重聚焦于相关校准实例,导致有效样本量下降,使标准共形p值过于保守以保障误差控制;而用于缓解此问题的平滑技术又引入条件方差,可能掩盖异常。为此,我们提出一种连续推断松弛方法,通过连续加权核密度估计将局部适应与尾部分辨率解耦。该方法虽放弃有限样本精确性转为渐近有效性,但消除了蒙特卡洛变异性,并恢复了因离散化损失的统计功效。实验表明,本方法不仅在离散基线失效时重获检测能力,且在统计功效上优于标准方法,同时在实践中保持有效的边际误差控制。
原文摘要 · Abstract (English)
Standard conformal anomaly detection provides marginal finite-sample guarantees under the assumption of exchangeability . However, real-world data often exhibit distribution shifts, necessitating a weighted conformal approach to adapt to local non-stationarity. We show that this adaptation induces a critical trade-off between the minimum attainable p-value and its stability. As importance weights localize to relevant calibration instances, the effective sample size decreases. This can render standard conformal p-values overly conservative for effective error control, while the smoothing technique used to mitigate this issue introduces conditional variance, potentially masking anomalies. We propose a continuous inference relaxation that resolves this dilemma by decoupling local adaptation from tail resolution via continuous weighted kernel density estimation. While relaxing finite-sample exactness to asymptotic validity, our method eliminates Monte Carlo variability and recovers the statistical power lost to discretization. Empirical evaluations confirm that our approach not only restores detection capabilities where discrete baselines yield zero discoveries, but outperforms standard methods in statistical power while maintaining valid marginal error control in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。