用集合先验提升半监督层次聚类的树结构准确性。
Semi-Supervised Hyperbolic Hierarchical Clustering with Set-Level Structural Priors

- 以集合为单位建模层次结构,从局部约束推导出整体子树先验。
- 在11个数据集上相比基线提升标签一致性,且优化了相似性质量。
- 适合需要精确层次结构的场景,如生物分类、文档组织。
半监督层次聚类旨在学习与数据模式和用户提供的监督一致的树结构。监督通常以叶级关系形式给出,如成对的必须链接/不能链接约束或三元组的必须先于约束。尽管这些监督有助于调节局部样本关系,但无法直接指示哪些样本应形成连贯的子树,导致学习到的非叶结构可能偏离真实标签所期望的层级组织。为解决此问题,我们提出一种带有集合级结构先验的半监督双曲层次聚类方法。主要贡献在于引入集合作为层次学习的基本建模单元。每个集合表示预期在子树内保持一致的样本集合,由叶级监督与学习到的一致相似性结构共同推导而来。这些集合作为软结构先验,指导非叶层级的形成,超越局部叶级关系。具体而言,首先学习一致约束的嵌入以获得可靠的集合划分,然后构建由约束诱导的集合并估计集合间相似性以形成集合级结构先验。最后,将这些先验融入双曲层次目标函数中,实现连续树优化。在11个基准数据集上的实验及消融研究显示,所提方法在标签一致性方面持续优于代表性层次聚类基线,同时提升了基于相似性的树质量。
原文摘要 · Abstract (English)
Semi-supervised hierarchical clustering aims to learn a tree structure consistent with data patterns and user-provided supervision. Supervision is usually given as leaf-level relations, such as pairwise must-link/cannot-link constraints or triplet-wise must-link-before constraints. Although useful for regulating local sample relations, such supervision does not directly indicate which samples should form coherent subtrees. Consequently, the non-leaf structure of the learned tree may deviate from the hierarchical organization preferred by ground-truth labels. To address this limitation, we propose a semi-supervised hyperbolic hierarchical clustering method with set-level structural priors. The main contribution is to introduce sets as basic modeling units for hierarchy learning. Each set denotes samples expected to cohere within a subtree and is induced from leaf-level supervision together with a learned constraint-consistent similarity structure. These sets act as soft structural priors for subtree-level supervision, allowing supervision to guide non-leaf hierarchy formation beyond local leaf-level relations. Specifically, we first learn constraint-consistent embeddings to obtain a reliable set partition, then construct constraint-induced sets and estimate inter-set similarities to form set-level structural priors. Finally, these priors are incorporated into a hyperbolic hierarchy objective for continuous tree optimization. Experiments on eleven benchmark datasets and ablation studies show that the proposed method consistently improves label consistency over representative hierarchical clustering baselines while also enhancing similarity-based tree quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。