构建包含2446个数据集的大规模表格异常检测基准,推动方法评估标准化。
MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection
- 构建三个分量:真实语义异常、统计异常与合成数据集,共2446个
- 涵盖790+真实异常数据集,支持跨领域方法对比与统计验证
- 提供标准划分与在线排行榜,适合研究者和工业界验证模型性能
高质量基准对科学进步和方法选择至关重要。当前表格异常检测(OD)基准仍受限,主流的AdBench仅含57个数据集,规模小、多样性不足。本文提出MacrOData,一个大规模表格OD基准套件,包含三个精心构建的组件:OddBench(790个含真实语义异常的数据集)、OvrBench(856个含真实统计异常的数据集)和SynBench(800个覆盖多种数据先验与异常类型的人工生成数据集)。该基准支持全面且统计稳健的评估。所有数据集均提供标准训练/测试划分,公开/私有分区及保留测试标签的线上排行榜。同时标注语义元数据。我们对经典、深度与基础模型在多组超参数下进行广泛实验,报告详细实证结果、实用建议及个体性能参考。全部2446个数据集已开源,配套排行榜部署于https://huggingface.co/MacrOData-CMU。
原文摘要 · Abstract (English)
Quality benchmarks are essential for fairly and accurately tracking scientific progress and enabling practitioners to make informed methodological choices. Outlier detection (OD) on tabular data underpins numerous real-world applications, yet existing OD benchmarks remain limited. The prominent OD benchmark AdBench is the de facto standard in the literature, yet comprises only 57 datasets. In addition to other shortcomings discussed in this work, its small scale severely restricts diversity and statistical power. We introduce MacrOData, a large-scale benchmark suite for tabular OD comprising three carefully curated components: OddBench, with 790 datasets containing real-world semantic anomalies; OvrBench, with 856 datasets featuring real-world statistical outliers; and SynBench, with 800 synthetically generated datasets spanning diverse data priors and outlier archetypes. Owing to its scale and diversity, MacrOData enables comprehensive and statistically robust evaluation of tabular OD methods. Our benchmarks further satisfy several key desiderata: We provide standardized train/test splits for all datasets, public/private benchmark partitions with held-out test labels for the latter reserved toward an online leaderboard, and annotate our datasets with semantic metadata. We conduct extensive experiments across all benchmarks, evaluating a broad range of OD methods comprising classical, deep, and foundation models, over diverse hyperparameter configurations. We report detailed empirical findings, practical guidelines, as well as individual performances as references for future research. All benchmarks containing 2,446 datasets combined are open-sourced, along with a publicly accessible leaderboard hosted at https://huggingface.co/MacrOData-CMU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。