提出新方法统一处理数值与类别属性聚类
New Approach to Clustering Random Attributes
- 将类别属性编码为数值形式,实现与数值属性统一聚类
- 支持数值属性、编码后类别属性及混合属性的聚类
- 方法通用性强,适用于多种数据集类型
本文提出一种新的相似性分析方法,进而设计了一种用于聚类不同类型随机属性(包括数值型和类别型)的新算法。为使类别属性可聚类,需将其值进行恰当编码,转化为数值形式。仅数值属性可进行因子分析,从而基于其与因子的相似性进行聚类。所提方法在多个样本数据集上进行了测试,结果表明该方法具有普遍适用性:既能对数值属性聚类,也能对经数值编码的类别属性聚类,还可实现数值属性与编码后类别属性的联合聚类。
原文摘要 · Abstract (English)
This paper proposes a new method for similarity analysis and, consequently, a new algorithm for clustering different types of random attributes, both numerical and nominal. However, in order for nominal attributes to be clustered, their values must be properly encoded. In the encoding process, nominal attributes obtain a new representation in numerical form. Only the numeric attributes can be subjected to factor analysis, which allows them to be clustered in terms of their similarity to factors. The proposed method was tested for several sample datasets. It was found that the proposed method is universal. On the one hand, the method allows clustering of numerical attributes. On the other hand, it provides the ability to cluster nominal attributes. It also allows simultaneous clustering of numerical attributes and numerically encoded nominal attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。