提出统一度量名义与有序属性距离的新方法,提升类别数据聚类效果。
Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes

- 基于图视角建模名义与有序属性的内在差异与关联
- 联合学习属性距离权重与数据聚类结果,避免次优解
- 在多个数据集上优于现有方法,尤其适合混合类型数据
类别数据聚类的成功高度依赖于衡量对象间差异的度量方式。然而,现有方法通常对名义属性和有序属性采用相同处理方式,忽略了有序属性间的相对顺序信息。此外,名义与有序属性之间存在相互依赖关系,值得深入探索以揭示差异性。本文从图论视角出发,研究两类属性值的内在差异与联系,提出一种新的统一距离度量方法,可同时保留有序属性的顺序关系。进一步提出新聚类算法,将属性距离权重学习与数据划分整合为单一学习范式,避免分步优化导致的次优解。实验表明,该方法在多个基准数据集上显著优于现有方法。
原文摘要 · Abstract (English)
The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical subtypes, i.e. nominal and ordinal attributes, in the same way when calculating the dissimilarity without considering the relative order information of the ordinal values. Moreover, there would exist interdependence among the nominal and ordinal attributes, which is worth exploring for indicating the dissimilarity. This paper will therefore study the intrinsic difference and connection of nominal and ordinal attribute values from a perspective akin to the graph. Accordingly, we propose a novel distance metric to measure the intra-attribute distances of nominal and ordinal attributes in a unified way, meanwhile preserving the order relationship among ordinal values. Subsequently, we propose a new clustering algorithm to make the learning of intra-attribute distance weights and partitions of data objects into a single learning paradigm rather than two separate steps, whereby circumventing a suboptimal solution. Experiments show the efficacy of the proposed algorithm in comparison with the existing counterparts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。