构建近3000个真实数据风格的合成聚类测试集,提升算法评估真实性。
ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

- 从200多个真实数据集衍生出灵活非参数合成数据
- 生成数据规模普遍大于原始数据,保留真实复杂性
- 配套R包开源,支持大规模聚类方法基准测试
尽管已有常用测试平台用于评估聚类方法性能,但大规模基准测试通常局限于较简单的模拟设置。本文介绍近3000个合成数据集的生成与整理工作,这些数据集源自200多个公开的真实数据集,多数来自实际应用场景。通过对每个原始数据集拟合灵活的非参数分布,我们保留了真实数据中难以在标准模拟中复现的细微特征,同时生成的数据集规模有时远超原始数据集。所有合成数据集及配套R包已开放下载,地址为https://github.com/DavidHofmeyr/ClusBench。
原文摘要 · Abstract (English)
Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relatively simplistic simulation set-ups. Here we describe the production and curation of close to 3000 synthetic data sets, derived from more than 200 publicly available data sets; the majority of which arose from real-world applications. By fitting a flexible non-parametric distribution to each base data set we are able to retain much of the nuance in real-world data which is difficult to reproduce in standard simulations, while also producing data sets whose sizes are sometimes substantially greater than the data sets from which they are derived. The synthetic data sets, plus an accompanying R package, are available for download from https://github.com/DavidHofmeyr/ClusBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。