基于坐标符号分区的高维数据聚类新方法,兼顾稳定性与拓扑洞察。
Hyperoctant Search Clustering: A Method for Clustering Data in High-Dimensional Hyperspheres
- 按坐标符号划分超象限,用图结构连接邻近区域进行聚类
- 在主题检测任务中对超参数变化更稳定,且能直接给出数据分布的超象限数
- 适合需要理解高维数据拓扑结构的研究者使用
高维数据聚类在人工智能、机器学习和模式识别中日益重要。本文提出一种基于组合拓扑方法的新聚类技术,该方法针对由坐标符号定义的空间区域(超象限)进行处理。在高维空间中,该方法通常可大幅缩减数据集规模,同时保留足够的拓扑特征。基于密度准则,方法通过构建图结构实现聚类:图的顶点代表超象限,边连接在莱文斯坦距离下相邻的超象限。该方法称为超象限搜索聚类(HyperOctant Search Clustering)。我们证明了该方法的部分数学性质。为评估性能,选择主题检测这一文本挖掘中的关键任务进行实验。结果表明,该方法在主要超参数变化下更具稳定性,且不仅是聚类工具,还能从拓扑角度探索数据集,直接提供包含数据点的超象限数量。我们还讨论了该方法与其他研究领域的潜在关联。
原文摘要 · Abstract (English)
Clustering of high-dimensional data sets is a growing need in artificial intelligence, machine learning and pattern recognition. In this paper, we propose a new clustering method based on a combinatorial-topological approach applied to regions of space defined by signs of coordinates (hyperoctants). In high-dimensional spaces, this approach often reduces the size of the dataset while preserving sufficient topological features. According to a density criterion, the method builds clusters of data points based on the partitioning of a graph, whose vertices represent hyperoctants, and whose edges connect neighboring hyperoctants under the Levenshtein distance. We call this method HyperOctant Search Clustering. We prove some mathematical properties of the method. In order to as assess its performance, we choose the application of topic detection, which is an important task in text mining. Our results suggest that our method is more stable under variations of the main hyperparameter, and remarkably, it is not only a clustering method, but also a tool to explore the dataset from a topological perspective, as it directly provides information about the number of hyperoctants where there are data points. We also discuss the possible connections between our clustering method and other research fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。