arXiv:2501.18028cs.LG2025-01被引 2

用新型距离度量提升KNN和K-means在噪声下的鲁棒性

KNN and K-means in Gini Prametric Spaces

  • 基于值与排序信息设计新距离度量,增强抗噪能力
  • 在16个UCI数据集上表现优于传统方法,收敛性可证明
  • 适合处理含异常值的实测数据,适用于稳健建模场景

本文提出基于吉尼参数空间的K均值与K近邻算法改进方法,替代传统度量空间。吉尼参数度量融合数值与排序信息,对噪声和异常值更具鲁棒性。主要贡献包括:(1) 提出一种同时捕捉数值距离与排序信息的吉尼参数度量;(2) 设计可证明收敛的吉尼K均值算法,对噪声数据具有强适应性;(3) 提出吉尼KNN方法,在噪声环境下性能媲美哈桑纳特等先进距离度量。在16个UCI数据集上的实验表明,该方法在聚类与分类任务中兼具优越性能与高效性。本工作为机器学习与统计分析中的排序型参数度量开辟新方向。

原文摘要 · Abstract (English)

This paper introduces enhancements to the K-means and K-nearest neighbors (KNN) algorithms based on the concept of Gini prametric spaces, instead of traditional metric spaces. Unlike standard distance metrics, Gini prametrics incorporate both value-based and rank-based measures, offering robustness to noise and outliers. The main contributions include: (1) a Gini prametric that captures rank information alongside value distances; (2) a Gini K-means algorithm that is provably convergent and resilient to noisy data; and (3) a Gini KNN method that performs competitively with state-of-the-art approaches like Hassanat's distance in noisy environments. Experimental evaluations on 16 UCI datasets demonstrate the superior performance and efficiency of the Gini-based algorithms in clustering and classification tasks. This work opens new directions for rank-based prametrics in machine learning and statistical analysis.

聚类距离度量鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。