统一学习异构属性距离度量,提升混合数据聚类效果
Learning Unified Distance Metric for Heterogeneous Attribute Data Clustering
- 将数值与类别属性映射到统一可学习空间,实现联合度量学习
- 在多个真实数据集上优于主流方法,聚类准确率显著提升
- 无需调参、收敛有保证,适合不同簇数的自适应聚类任务
由数值和类别属性组成的混合数据在实际聚类任务中常见。与数值属性在明确的欧氏空间中表示概念间趋势(如高低温度)不同,类别属性值是不同概念(如职业)嵌入于隐式空间。同时利用这两种差异巨大的信息是不可避免但极具挑战的问题。现有方法多将异构属性统一编码为一类,或定义统一度量,却未揭示其内在关联。本文研究各类属性间的联系,提出新颖的异构属性重建与表征(HARR)学习范式,用于聚类分析。该范式将异构属性转换为同质状态以进行距离度量学习,并将学习过程与聚类结合,自动适应不同聚类任务。不同于多数工作直接使用预定义度量或学习属性权重搜索子空间,本文提出将每项属性值投影至统一的可学习多空间,更精细地表示并学习类别数据的距离度量。HARR无需参数、收敛性有保障,能更有效地自适应不同簇数 $k$。大量实验表明其在准确性和效率上均具优势。
原文摘要 · Abstract (English)
Datasets composed of numerical and categorical attributes (also called mixed data hereinafter) are common in real clustering tasks. Differing from numerical attributes that indicate tendencies between two concepts (e.g., high and low temperature) with their values in well-defined Euclidean distance space, categorical attribute values are different concepts (e.g., different occupations) embedded in an implicit space. Simultaneously exploiting these two very different types of information is an unavoidable but challenging problem, and most advanced attempts either encode the heterogeneous numerical and categorical attributes into one type, or define a unified metric for them for mixed data clustering, leaving their inherent connection unrevealed. This paper, therefore, studies the connection among any-type of attributes and proposes a novel Heterogeneous Attribute Reconstruction and Representation (HARR) learning paradigm accordingly for cluster analysis. The paradigm transforms heterogeneous attributes into a homogeneous status for distance metric learning, and integrates the learning with clustering to automatically adapt the metric to different clustering tasks. Differing from most existing works that directly adopt defined distance metrics or learn attribute weights to search clusters in a subspace. We propose to project the values of each attribute into unified learnable multiple spaces to more finely represent and learn the distance metric for categorical data. HARR is parameter-free, convergence-guaranteed, and can more effectively self-adapt to different sought number of clusters $k$. Extensive experiments illustrate its superiority in terms of accuracy and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。