arXiv:2508.18474cs.LGcs.AI2025-08

动态奖励调节提升时间序列异常检测精度

DRTA: Dynamic Reward Scaling for Reinforcement Learning in Time Series Anomaly Detection

  • 用VAE重构误差与分类奖励动态调参,平衡探索与利用
  • 在Yahoo A1/A2数据集上优于现有无监督/半监督方法
  • 适合标签稀缺且需高精度的工业监控场景

时间序列异常检测在金融、医疗、传感器网络和工业监控中至关重要。传统方法常受限于标注数据少、误报率高,且难以泛化到新型异常。为此,我们提出基于强化学习的DRTA框架,融合动态奖励设计、变分自编码器(VAE)与主动学习。该方法通过自适应机制动态调节VAE重构误差与分类奖励的影响,实现探索与利用的平衡,从而在低标签条件下有效检测异常,同时保持高精确率与召回率。在Yahoo A1和Yahoo A2基准数据集上的实验表明,该方法持续优于当前最优的无监督与半监督方法。结果证明,该框架是可扩展且高效的现实异常检测解决方案。

原文摘要 · Abstract (English)

Anomaly detection in time series data is important for applications in finance, healthcare, sensor networks, and industrial monitoring. Traditional methods usually struggle with limited labeled data, high false-positive rates, and difficulty generalizing to novel anomaly types. To overcome these challenges, we propose a reinforcement learning-based framework that integrates dynamic reward shaping, Variational Autoencoder (VAE), and active learning, called DRTA. Our method uses an adaptive reward mechanism that balances exploration and exploitation by dynamically scaling the effect of VAE-based reconstruction error and classification rewards. This approach enables the agent to detect anomalies effectively in low-label systems while maintaining high precision and recall. Our experimental results on the Yahoo A1 and Yahoo A2 benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art unsupervised and semi-supervised approaches. These findings show that our framework is a scalable and efficient solution for real-world anomaly detection tasks.

异常检测强化学习时间序列动态奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。