提出一种能处理重叠与离群点的聚类比较方法。
A Pragmatic Method for Comparing Clusterings with Overlaps and Outliers
- 设计了一种兼顾重叠与离群点的聚类相似性度量。
- 实验验证该方法不受常见偏差影响,结果更可靠。
- 适合存在重叠或离群点的真实数据聚类评估。
聚类算法是无监督数据科学的核心组成部分,其外部评估需要一种将检测到的聚类与真实聚类进行比较的方法。在一般情况下,检测到的聚类和真实聚类可能包含离群点(不属于任何簇的样本)、重叠簇(样本可属于多个簇),或两者兼具,但目前尚无有效的比较方法。本文提出一种实用的相似性度量,用于比较含重叠与离群点的聚类,并证明该度量具备若干理想性质,实验进一步证实其不受其他聚类比较度量中常见的多种偏差影响。
原文摘要 · Abstract (English)
Clustering algorithms are an essential part of the unsupervised data science ecosystem, and extrinsic evaluation of clustering algorithms requires a method for comparing the detected clustering to a ground truth clustering. In a general setting, the detected and ground truth clusterings may have outliers (objects belonging to no cluster), overlapping clusters (objects may belong to more than one cluster), or both, but methods for comparing these clusterings are currently undeveloped. In this note, we define a pragmatic similarity measure for comparing clusterings with overlaps and outliers, show that it has several desirable properties, and experimentally confirm that it is not subject to several common biases afflicting other clustering comparison measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。