用多轮对话增强大模型的时序异常检测能力,提升解释性与泛化性。
ChatAD: Reasoning-Enhanced Time-Series Anomaly Detection with Multi-Turn Instruction Evolution
- 设计多智能体演化框架TSEvol,支持多轮指令迭代优化
- 在7个数据集上实现最高34.5%准确率提升,误报减少37.42%
- 适合作为需解释性的工业异常检测系统技术参考
基于大语言模型的时序异常检测(AD)能增强对异常行为的理解与解释能力。现有方法存在推理能力不足、多轮对话支持弱、泛化性差等问题。为此,本文提出基于多智能体的时序演化算法TSEvol;构建包含2万条样本的AD多轮对话数据集TSEData-20K,推出ChatAD-Llama3-8B、Qwen2.5-7B和Mistral-7B三款对话式检测模型;引入时序凯恩曼-特韦斯基优化(TKTO)以提升跨任务泛化能力;并建立基于大模型的评估基准LLADBench,涵盖7个数据集与多种任务。实验表明,三款ChatAD模型在准确率上最高提升34.50%,F1值提升34.71%,误报率降低37.42%。通过TKTO优化后,模型在分类、预测与填补任务中均表现出色。
原文摘要 · Abstract (English)
LLM-driven Anomaly Detection (AD) helps enhance the understanding and explanatory abilities of anomalous behaviors in Time Series (TS). Existing methods face challenges of inadequate reasoning ability, deficient multi-turn dialogue capability, and narrow generalization. To this end, we 1) propose a multi-agent-based TS Evolution algorithm named TSEvol. On top of it, we 2) introduce the AD reasoning and multi-turn dialogue Dataset TSEData-20K and contribute the Chatbot family for AD, including ChatAD-Llama3-8B, Qwen2.5-7B, and Mistral-7B. Furthermore, 3) we propose the TS Kahneman-Tversky Optimization (TKTO) to enhance ChatAD's cross-task generalization capability. Lastly, 4) we propose a LLM-driven Learning-based AD Benchmark LLADBench to evaluate the performance of ChatAD and nine baselines across seven datasets and tasks. Our three ChatAD models achieve substantial gains, up to 34.50% in accuracy, 34.71% in F1, and a 37.42% reduction in false positives. Besides, via KTKO, our optimized ChatAD achieves competitive performance in reasoning and cross-task generalization on classification, forecasting, and imputation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。