通过特征重要性重标定,提升噪声数据中聚类评估的准确性。
Improving clustering quality evaluation in noisy Gaussian mixtures
- 基于特征离散度动态调整贡献,抑制噪声特征干扰。
- 在合成与真实数据上显著提高评估指标与真实聚类质量的相关性。
- 特别适合高维、含噪或重叠簇的无监督学习场景。
聚类是机器学习和数据分析中的经典技术,广泛应用于多个领域。当缺乏外部真实标签时,聚类有效性指标(如平均轮廓宽、Calinski-Harabasz、Davies-Bouldin)用于评估聚类质量。然而,这些指标易受特征相关性差异影响,在高维或噪声数据中可能产生不可靠结果。本文提出理论严谨的特征重要性重标定(FIR)方法,根据特征分散程度调整其贡献,削弱噪声特征影响,增强聚类紧凑性与分离性,使评估更贴近真实情况。在不同配置的合成数据集及真实世界数据案例研究中,FIR均显著提升有效性指标与真实聚类质量之间的相关性,尤其在存在噪声或无关特征时表现更优。结果表明,FIR增强了聚类评估的鲁棒性,降低了跨数据集的性能波动,即使在簇间重叠严重时仍有效。这表明FIR可作为聚类评估的有效增强工具,适用于缺乏标注数据的无监督学习任务。
原文摘要 · Abstract (English)
Clustering is a well-established technique in machine learning and data analysis, widely used across various domains. Cluster validity indices, such as the Average Silhouette Width, Calinski-Harabasz, and Davies-Bouldin indices, play a crucial role in assessing clustering quality when external ground truth labels are unavailable. However, these measures can be affected by different degrees of feature relevance, potentially leading to unreliable evaluations in high-dimensional or noisy data sets. We introduce a theoretically grounded Feature Importance Rescaling (FIR) method that enhances the quality of clustering validation by adjusting feature contributions based on their dispersion. It attenuates noise features, clarifies clustering compactness and separation, and thereby aligns clustering validation more closely with the ground truth. Through extensive experiments on synthetic data sets under different configurations and a case study on real-world data, we demonstrate that FIR consistently improves the correlation between the values of cluster validity indices and the ground truth, particularly in settings with noisy or irrelevant features. The results show that FIR increases the robustness of clustering evaluation, reduces variability in performance across different data sets, and remains effective even when clusters exhibit significant overlap. These findings highlight the potential of FIR as a valuable enhancement of clustering validation, making it a practical tool for unsupervised learning tasks where labelled data is unavailable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。