arXiv:2409.03046cs.CL2024-09被引 2

用语言模型检测文本异常,新指标比传统方法更有效。

Oddballness: universal anomaly detection with language models

  • 提出新指标'oddballness'衡量令牌异常程度
  • 在无监督语法纠错任务中优于传统低概率检测
  • 适合无标注数据的异常检测场景

我们提出一种全新的无监督文本异常检测方法,利用语言模型生成的概率,但不关注低概率词元,而是引入本文提出的全新指标——oddballness,用于衡量给定词元相对于语言模型的‘奇怪程度’。在语法错误检测(文本异常检测的一个具体案例)任务中,实验表明,在完全无监督设置下,oddballness的表现优于仅依赖低概率事件的传统方法。

原文摘要 · Abstract (English)

We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new metric introduced in this paper: oddballness. Oddballness measures how ``strange'' a given token is according to the language model. We demonstrate in grammatical error detection tasks (a specific case of text anomaly detection) that oddballness is better than just considering low-likelihood events, if a totally unsupervised setup is assumed.

异常检测语言模型无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。