arXiv:2412.16713math.NAcs.LG2024-12被引 1

一种统一的数据自适应聚类算法,能自动发现数据结构并适配多种场景。

A Unifying Family of Data-Adaptive Partitioning Algorithms

  • 通过单一参数控制,统一k-means与k-子空间等经典方法
  • 无需专家知识即可自动识别数据结构与参数
  • 适用于图像分割、模型降阶、矩阵近似等多领域

聚类算法仍是分组和总结数据关键特征的有力工具,广泛应用于图像分割、降维、信号分析、模型降阶、数值分析等领域。为此,已发展出多种满足特定需求的方法。本文提出一个数据自适应划分算法族,统一了k-means、k-子空间等若干知名方法。该算法族由单一参数索引,采用统一最小化策略,易于使用与解释,并可高效处理大规模高维问题。此外,我们设计了一种自适应机制,能够(a)无需专家知识自动揭示数据结构与问题参数,(b)可增强现有其他方法。通过在子空间聚类、模型降阶和矩阵逼近等不同领域的应用演示,展示了其通用性与拓展科学边界的潜力。我们认为,该家族的参数化结构实现了算法间的协同效应,有望推动数据科学领域的新发展。

原文摘要 · Abstract (English)

Clustering algorithms remain valuable tools for grouping and summarizing the most important aspects of data. Example areas where this is the case include image segmentation, dimension reduction, signals analysis, model order reduction, numerical analysis, and others. As a consequence, many clustering approaches have been developed to satisfy the unique needs of each particular field. In this article, we present a family of data-adaptive partitioning algorithms that unifies several well-known methods (e.g., k-means and k-subspaces). Indexed by a single parameter and employing a common minimization strategy, the algorithms are easy to use and interpret, and scale well to large, high-dimensional problems. In addition, we develop an adaptive mechanism that (a) exhibits skill at automatically uncovering data structures and problem parameters without any expert knowledge and, (b) can be used to augment other existing methods. By demonstrating the performance of our methods on examples from disparate fields including subspace clustering, model order reduction, and matrix approximation, we hope to highlight their versatility and potential for extending the boundaries of existing scientific domains. We believe our family's parametrized structure represents a synergism of algorithms that will foster new developments and directions, not least within the data science community.

聚类算法数据自适应参数化模型多领域应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。