重建表格异常检测的语义上下文,提升模型理解与解释能力
ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
- 构建20个含结构化文本元数据的表格数据集,还原真实场景中的领域知识
- 引入零样本LLM框架,在不微调的情况下利用语义上下文提升检测效果
- 适合关注可解释性、领域知识融合的异常检测研究者使用
在表格异常检测中,文本语义常包含关键信号,因异常定义高度依赖领域上下文。现有基准仅提供原始数据点,缺乏特征描述、领域知识等丰富文本元数据,限制了模型对领域知识的利用。本文提出ReTabAD,通过恢复文本语义,推动上下文感知的表格异常检测研究。该基准包含20个精心构建的表格数据集,附带经典、深度学习及基于大模型的先进算法实现;同时提出一个无需任务特定训练的零样本大模型框架,建立强基线。实验表明,语义上下文能显著提升检测性能,并支持领域相关推理,增强可解释性。此工作为系统探索上下文感知异常检测提供了新范式。
原文摘要 · Abstract (English)
In tabular anomaly detection (AD), textual semantics often carry critical signals, as the definition of an anomaly is closely tied to domain-specific context. However, existing benchmarks provide only raw data points without semantic context, overlooking rich textual metadata such as feature descriptions and domain knowledge that experts rely on in practice. This limitation restricts research flexibility and prevents models from fully leveraging domain knowledge for detection. ReTabAD addresses this gap by restoring textual semantics to enable context-aware tabular AD research. We provide (1) 20 carefully curated tabular datasets enriched with structured textual metadata, together with implementations of state-of-the-art AD algorithms including classical, deep learning, and LLM-based approaches, and (2) a zero-shot LLM framework that leverages semantic context without task-specific training, establishing a strong baseline for future research. Furthermore, this work provides insights into the role and utility of textual metadata in AD through experiments and analysis. Results show that semantic context improves detection performance and enhances interpretability by supporting domain-aware reasoning. These findings establish ReTabAD as a benchmark for systematic exploration of context-aware AD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。