arXiv:2605.30700cs.CVcs.LG2026-05

用数学形态学提升聚类与分类,更好捕捉数据形状和密度特征。

Mathematical Morphology in Machine Learning

  • 基于形态学重建的快速聚类,能保持簇形貌与密度分布。
  • 新距离度量在33个数据集上26次表现优于平均,9次最优。
  • 首次融合形状、密度与分形信息构建分类器,适合复杂结构数据。

本文将成熟的视觉计算理论——数学形态学引入机器学习,以挖掘标准方法常忽略的形状与密度信息。提出一种基于形态学重建的快速聚类算法,能准确保持簇的形状与密度分布,具备内在最大簇识别能力、无需额外代价的噪声清除功能,以及由结构元素调控的多样生长模式。此外,设计了一种结合Minkowski与Chebyshev距离的新度量,在$Z^2$离散邻域迭代中速度约为曼哈顿距离的1.3倍,比欧几里得距离快329.5倍。在33个UCI数据集上的k-NN分类实验中,该度量在26次中超过平均性能,9次取得最佳准确率。最后,提出新型形态学分类器,首次同时建模数据集的形状、密度与分形特性。

原文摘要 · Abstract (English)

This work introduces mathematical morphology-an established visual computing theory-into machine learning to exploit shape and density aspects often overlooked by standard techniques. We propose a fast clustering algorithm based on morphological reconstruction that accurately preserves cluster shapes and density. This scheme offers unique features: an intrinsic sense of maximal clusters, cost-free noise removal, and diverse growth patterns controlled by structuring elements.Additionally, we propose a novel distance metric combining Minkowski and Chebyshev distances, highly efficient for morphological dilations. In $Z^2$ discrete neighbourhood iterations, it is roughly 1.3 times faster than Manhattan and 329.5 times faster than Euclidean distances. When evaluated using a k-Nearest Neighbours (k-NN) classifier across 33 UCI datasets against 14 other distances, our metric achieved above-average accuracies most frequently (26 of 33 cases) and the best overall accuracy in 9 cases.Finally, we introduce novel morphological classifiers. Unlike current literature, this proposal uniquely models shape, density, and fractal information in datasets.

数学形态学聚类算法距离度量分类器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。