arXiv:2410.16124cs.LGstat.ML2024-10被引 1

用合成数据集MNIST-Nd测试高维聚类效果,发现维度越高越难聚类。

MNIST-Nd: a set of naturalistic datasets to benchmark clustering across dimensions

  • 基于变分自编码器生成6个不同维度的合成数据集
  • 实验表明维度越高,聚类难度越大,Leiden算法最稳定
  • 适合研究高维数据聚类性能或算法鲁棒性的学者使用

随着记录技术进步,各科学领域涌现出大量大规模高维数据。尤其在生物学中,聚类常用于揭示数据结构,如细胞类型组织方式。然而,聚类在高维下表现不佳,且维度影响尚不明确,因现有基准数据集多为二维。本文提出MNIST-Nd,一组通过在MNIST上训练具有2至64个隐变量的混合变分自编码器生成的合成数据集,包含6个具有相似结构但维度不同的数据集,可分离维度对聚类的影响。初步基准测试显示,随维度增加,Leiden算法在聚类中表现最稳健。

原文摘要 · Abstract (English)

Driven by advances in recording technology, large-scale high-dimensional datasets have emerged across many scientific disciplines. Especially in biology, clustering is often used to gain insights into the structure of such datasets, for instance to understand the organization of different cell types. However, clustering is known to scale poorly to high dimensions, even though the exact impact of dimensionality is unclear as current benchmark datasets are mostly two-dimensional. Here we propose MNIST-Nd, a set of synthetic datasets that share a key property of real-world datasets, namely that individual samples are noisy and clusters do not perfectly separate. MNIST-Nd is obtained by training mixture variational autoencoders with 2 to 64 latent dimensions on MNIST, resulting in six datasets with comparable structure but varying dimensionality. It thus offers the chance to disentangle the impact of dimensionality on clustering. Preliminary common clustering algorithm benchmarks on MNIST-Nd suggest that Leiden is the most robust for growing dimensions.

聚类高维数据合成数据机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。