用语言模型检测文本异常,新指标比传统方法更有效。
Oddballness: universal anomaly detection with language models
- 提出新指标'oddballness'衡量令牌异常程度
- 在无监督语法纠错任务中优于传统低概率检测
- 适合无标注数据的异常检测场景
我们提出一种全新的无监督文本异常检测方法,利用语言模型生成的概率,但不关注低概率词元,而是引入本文提出的全新指标——oddballness,用于衡量给定词元相对于语言模型的‘奇怪程度’。在语法错误检测(文本异常检测的一个具体案例)任务中,实验表明,在完全无监督设置下,oddballness的表现优于仅依赖低概率事件的传统方法。
原文摘要 · Abstract (English)
We present a new method to detect anomalies in texts (in general: in sequences of any data), using language models, in a totally unsupervised manner. The method considers probabilities (likelihoods) generated by a language model, but instead of focusing on low-likelihood tokens, it considers a new metric introduced in this paper: oddballness. Oddballness measures how ``strange'' a given token is according to the language model. We demonstrate in grammatical error detection tasks (a specific case of text anomaly detection) that oddballness is better than just considering low-likelihood events, if a totally unsupervised setup is assumed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。