arXiv:2607.11969stat.MLcs.LG2026-07被引 1

测试发现多数异常检测评估指标在多次尝试中会失效,仅单次运行结果可信。

Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection

论文配图:Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection
图 1 · 摘自论文原文
  • 通过对抗性测试验证新评估指标在真实数据集上的鲁棒性
  • 单次运行下随机检测得分不超过最优的90%,但多次尝试后指标严重虚高
  • 推荐使用单次运行结果或披露尝试次数,优先选PR类指标

点调整(PA)作为时间序列异常检测(TSAD)的标准评分方式,被证实可给随机异常分数赋予接近完美的F1值。为此,学界提出一系列替代指标(PA%K、基于范围的精确率/召回率、归属精确率/召回率、体积-曲线下面积,VUS,ROC/PR)。我们独立且对抗性地测试这些指标在真实基准上的抗随机探测能力,发现结果完全取决于一个被忽视的变量:N,即对手报告最佳结果所进行的随机尝试次数。在单次诚实运行(N=1)下,六大数据集(UCR、SMD、SMAP、MSL、NAB、PSM)上无一替代指标可被攻破:随机检测在最多11%的序列上达到真实最优检测器得分的90%(归属-F1),5%(ROC族),2%(PR族和PA%K)。但在最佳-之一(best-of-N)报告下,指标表现急剧分化:归属-F1与所有ROC类指标显著膨胀,归属在N=3时即进入可被攻破状态(25%序列),至N=41时达0.98;ROC族在N=9-11时突破阈值;而PR类指标与PA%K在各N值下基本持平,接近异常比例(唯一例外为大N下的NAB)。成对检验显示,VUS-ROC在131个序列上被抬高,而其孪生指标VUS-PR未被影响,反之不成立。这一分裂源于极端类别不平衡下AUC的顺序统计特性(随机PR-AUC被锚定于异常比例);归属指标的膨胀则源于其单次运行的极端宽容性(即使在N=1时已脆弱)。我们发布了一个可pip安装的压力测试工具包,并建议报告单次运行结果或披露N值,优先选择抗最佳-之一膨胀的PR类指标。

原文摘要 · Abstract (English)

Point-adjustment (PA), for years the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al. (2022) to award near-perfect F1 to random anomaly scores. The field adopted a suite of replacement metrics (PA%K, range-based precision/recall, affiliation precision/recall, and Volume-Under-the-Surface, VUS, ROC/PR). We ask, independently and adversarially, whether these resist no-skill detectors on real benchmarks, and find the answer turns entirely on one overlooked variable: N, the number of random attempts an adversary reports the best of. Under a single honest run (N=1), not one replacement metric is gameable on any of six benchmarks (UCR, SMD, SMAP, MSL, NAB, PSM): a random detector reaches 90% of the best real detector's score on at most 11% of series for affiliation-F1, 5% for the ROC family, and 2% for the PR-based metrics and PA%K. But under best-of-N reporting, the seed-shopping endemic to ML, the metrics split sharply. affiliation-F1 and every ROC-based metric inflate steeply, affiliation crossing gameable (25% of series) by N=3 and reaching 0.98 at the full pool (N=41), the ROC family crossing by N=9-11; the PR-based metrics and PA%K stay near-flat at every N, floored near the anomaly prevalence (the lone exception is NAB at large N). A paired test finds VUS-ROC inflated on 131 series where its sibling VUS-PR is not, and never the reverse. The ROC-vs-PR split follows from the order-statistic behaviour of AUC under extreme class imbalance (a random PR-AUC is floored at prevalence); affiliation inflates by a second route, its extreme single-run leniency (already fragile at N=1). We release a pip-installable stress-test harness, and recommend reporting single-run scores or disclosing N and preferring PR-based metrics, which resist best-of-N inflation on nearly every benchmark.

异常检测评估指标对抗测试时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。