新评估指标CCE同时衡量预测置信度与一致性,更稳定高效。
CCE: Confidence-Consistency Evaluation for Time Series Anomaly Detection
- 用贝叶斯估计量化异常分数不确定性,构建全局与事件级评分
- 理论证明有界性、抗扰动鲁棒性,计算复杂度仅O(n)
- 配套基准RankEval实现指标客观对比,开源可复现
时间序列异常检测的评估指标对模型评价至关重要。然而现有指标存在判别力不足、超参数敏感、对扰动脆弱及计算开销高等问题。本文提出置信度-一致性评估(CCE),通过贝叶斯估计量化异常分数的不确定性,构建全局与事件级的置信度和一致性评分,形成简洁的综合指标。理论上和实验上均证明CCE具有严格有界性、对分数扰动的Lipschitz鲁棒性,且时间复杂度为O(n)。此外,我们建立了RankEval基准,首次实现评估指标排名能力的标准化、可复现比较。CCE与RankEval均已开源。
原文摘要 · Abstract (English)
Time Series Anomaly Detection metrics serve as crucial tools for model evaluation. However, existing metrics suffer from several limitations: insufficient discriminative power, strong hyperparameter dependency, sensitivity to perturbations, and high computational overhead. This paper introduces Confidence-Consistency Evaluation (CCE), a novel evaluation metric that simultaneously measures prediction confidence and uncertainty consistency. By employing Bayesian estimation to quantify the uncertainty of anomaly scores, we construct both global and event-level confidence and consistency scores for model predictions, resulting in a concise CCE metric. Theoretically and experimentally, we demonstrate that CCE possesses strict boundedness, Lipschitz robustness against score perturbations, and linear time complexity $\mathcal{O}(n)$. Furthermore, we establish RankEval, a benchmark for comparing the ranking capabilities of various metrics. RankEval represents the first standardized and reproducible evaluation pipeline that enables objective comparison of evaluation metrics. Both CCE and RankEval implementations are fully open-source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。