arXiv:2412.11142cs.CLcs.AI2024-12ACL被引 43

首个评测大模型在异常检测中表现的基准,验证其零样本检测能力。

AD-LLM: Benchmarking Large Language Models for Anomaly Detection

  • 用预训练知识实现无需微调的零样本异常检测
  • 设计合理的数据增强方法可提升模型性能
  • 适合关注大模型在无监督异常检测应用的研究者

异常检测(AD)是机器学习中的重要任务,广泛应用于欺诈检测、医疗诊断和工业监控等领域。在自然语言处理(NLP)中,AD可用于识别垃圾信息、虚假内容和异常用户行为。尽管大语言模型(LLMs)在文本生成和摘要等任务上表现卓越,但其在异常检测中的潜力尚未充分探索。本文提出AD-LLM,首个针对NLP异常检测的大语言模型基准评估体系。研究涵盖三个关键任务:(i) 零样本检测,利用LLMs的预训练知识实现无需特定任务训练的异常检测;(ii) 数据增强,生成合成数据与类别描述以提升模型性能;(iii) 模型选择,借助LLMs推荐适用于特定数据集的无监督异常检测模型。通过多个数据集的实验发现,LLMs在零样本检测中表现良好,精心设计的数据增强方法有效,而基于数据集特征的模型选择解释仍具挑战性。据此,本文提出六个未来研究方向。

原文摘要 · Abstract (English)

Anomaly detection (AD) is an important machine learning task with many real-world uses, including fraud detection, medical diagnosis, and industrial monitoring. Within natural language processing (NLP), AD helps detect issues like spam, misinformation, and unusual user activity. Although large language models (LLMs) have had a strong impact on tasks such as text generation and summarization, their potential in AD has not been studied enough. This paper introduces AD-LLM, the first benchmark that evaluates how LLMs can help with NLP anomaly detection. We examine three key tasks: (i) zero-shot detection, using LLMs' pre-trained knowledge to perform AD without tasks-specific training; (ii) data augmentation, generating synthetic data and category descriptions to improve AD models; and (iii) model selection, using LLMs to suggest unsupervised AD models. Through experiments with different datasets, we find that LLMs can work well in zero-shot AD, that carefully designed augmentation methods are useful, and that explaining model selection for specific datasets remains challenging. Based on these results, we outline six future research directions on LLMs for AD.

异常检测大模型零样本NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。