用大模型增强强化学习,实现少标注下的高精度时序异常检测
LLM-Enhanced Reinforcement Learning for Time Series Anomaly Detection
- 结合大模型语义奖励与强化学习,引导智能体探索
- 在有限标注下达到顶尖检测准确率,优于现有方法
- 适合标注成本高的金融、医疗等真实场景
时序异常检测在金融、医疗、传感器网络和工业监控中至关重要。但该任务常面临标签稀疏、时序模式复杂和专家标注成本高等问题。本文提出统一框架,融合大语言模型(LLM)生成的潜在函数用于奖励塑造,结合强化学习(RL)、变分自编码器(VAE)增强的动态奖励缩放,以及带标签传播的主动学习。基于LSTM的强化学习智能体利用LLM提供的语义奖励指导探索,同时通过VAE重构误差引入无监督异常信号。主动学习选择最不确定样本,标签传播高效扩展标注数据。在Yahoo-A1和SMD基准上的实验表明,该方法在有限标注预算下达到当前最优检测精度,且在数据受限场景中表现良好。研究展示了将大模型与强化学习及先进无监督技术结合,在真实应用中实现鲁棒、可扩展异常检测的巨大潜力。
原文摘要 · Abstract (English)
Detecting anomalies in time series data is crucial for finance, healthcare, sensor networks, and industrial monitoring applications. However, time series anomaly detection often suffers from sparse labels, complex temporal patterns, and costly expert annotation. We propose a unified framework that integrates Large Language Model (LLM)-based potential functions for reward shaping with Reinforcement Learning (RL), Variational Autoencoder (VAE)-enhanced dynamic reward scaling, and active learning with label propagation. An LSTM-based RL agent leverages LLM-derived semantic rewards to guide exploration, while VAE reconstruction errors add unsupervised anomaly signals. Active learning selects the most uncertain samples, and label propagation efficiently expands labeled data. Evaluations on Yahoo-A1 and SMD benchmarks demonstrate that our method achieves state-of-the-art detection accuracy under limited labeling budgets and operates effectively in data-constrained settings. This study highlights the promise of combining LLMs with RL and advanced unsupervised techniques for robust, scalable anomaly detection in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。