arXiv:2501.00107cs.LGcs.AI2025-01被引 16

用强化学习动态选模型,无监督检测用电异常更准。

An Unsupervised Anomaly Detection in Electricity Consumption Using Reinforcement Learning and Time Series Forest Based Framework

  • 用强化学习自动挑选最适合的异常检测模型
  • 合成数据上F1达0.989,优于多数现有方法
  • 适合处理多种异常类型,尤其缺标签场景

异常检测(AD)在时间序列应用中至关重要,因时间序列广泛用于真实场景。由于异常形态多样,精准定位极具挑战。以往研究多依赖特定假设,对某些异常类型敏感性不足。本文提出一种结合时间序列森林(TSF)与强化学习(RL)的无监督模型选择框架,动态选取最优异常检测技术。该方法无需依赖稀缺且昂贵的标注数据,实现高效异常检测。真实数据实验显示,所提方法在F1分数上优于所有其他模型。合成数据上,除KNN外,其性能超越所有其他模型,达到0.989的高F1值。此外,在合成数据上,该框架表现优于提示GPT-4作为异常检测器的结果。不同奖励函数测试表明,原始奖励函数取得最佳综合性能。在包含全局、局部及聚类异常的三个附加数据集上评估六种模型,结果表明各模型性能随异常类型变化显著。本框架在所有数据集上均保持高精度,展现出跨异常类型的卓越性能。

原文摘要 · Abstract (English)

Anomaly detection (AD) plays a crucial role in time series applications, primarily because time series data is employed across real-world scenarios. Detecting anomalies poses significant challenges since anomalies take diverse forms making them hard to pinpoint accurately. Previous research has explored different AD models, making specific assumptions with varying sensitivity toward particular anomaly types. To address this issue, we propose a novel model selection for unsupervised AD using a combination of time series forest (TSF) and reinforcement learning (RL) approaches that dynamically chooses an AD technique. Our approach allows for effective AD without explicitly depending on ground truth labels that are often scarce and expensive to obtain. Results from the real-time series dataset demonstrate that the proposed model selection approach outperforms all other AD models in terms of the F1 score metric. For the synthetic dataset, our proposed model surpasses all other AD models except for KNN, with an impressive F1 score of 0.989. The proposed model selection framework also exceeded the performance of GPT-4 when prompted to act as an anomaly detector on the synthetic dataset. Exploring different reward functions revealed that the original reward function in our proposed AD model selection approach yielded the best overall scores. We evaluated the performance of the six AD models on an additional three datasets, having global, local, and clustered anomalies respectively, showing that each AD model exhibited distinct performance depending on the type of anomalies. This emphasizes the significance of our proposed AD model selection framework, maintaining high performance across all datasets, and showcasing superior performance across different anomaly types.

异常检测强化学习时间序列无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。