构建首个嵌入式文本异常检测综合评测基准
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
- 整合多领域数据集与主流大模型嵌入
- 系统验证不同嵌入与检测方法的组合效果
- 为真实场景异常检测提供可复用评估框架
文本异常检测在识别垃圾信息、虚假内容和不当言论方面至关重要。尽管嵌入式方法被广泛采用,其在不同应用场景下的有效性与泛化能力仍缺乏系统研究。为此,我们提出TAD-Bench,一个综合性基准,用于系统评估基于嵌入的文本异常检测方法。该基准整合了多个跨领域的数据集,结合大语言模型的先进嵌入表示与多种异常检测算法。通过大量实验,分析嵌入与检测方法之间的相互作用,揭示其在不同任务中的优势、局限性及适用性。研究结果为构建更鲁棒、高效且通用的异常检测系统提供了新视角,适用于真实世界应用。
原文摘要 · Abstract (English)
Text anomaly detection is crucial for identifying spam, misinformation, and offensive language in natural language processing tasks. Despite the growing adoption of embedding-based methods, their effectiveness and generalizability across diverse application scenarios remain under-explored. To address this, we present TAD-Bench, a comprehensive benchmark designed to systematically evaluate embedding-based approaches for text anomaly detection. TAD-Bench integrates multiple datasets spanning different domains, combining state-of-the-art embeddings from large language models with a variety of anomaly detection algorithms. Through extensive experiments, we analyze the interplay between embeddings and detection methods, uncovering their strengths, weaknesses, and applicability to different tasks. These findings offer new perspectives on building more robust, efficient, and generalizable anomaly detection systems for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。