跨数据集聚类新方法,自动适配差异提升效果
Adaptive Transfer Clustering: A Unified Framework
- 基于偏差-方差分解,自适应融合主辅数据聚类
- 在高斯混合模型下理论证明性能最优
- 适用于多种统计模型,适合多源数据聚类
我们提出一种通用的迁移学习聚类框架,针对具有相同研究对象的主数据集和辅助数据集。两个数据集可能反映相似但不同的潜在分组结构。本文提出自适应迁移聚类(ATC)算法,通过优化估计的偏差-方差分解,在未知差异存在时自动利用共性。该方法适用于广义统计模型,包括高斯混合模型、随机块模型和潜类别模型。理论分析证明了在高斯混合模型下ATC的最优性,并显式量化了迁移带来的增益。大量模拟与真实数据实验验证了该方法在多种场景下的有效性。
原文摘要 · Abstract (English)
We propose a general transfer learning framework for clustering given a main dataset and an auxiliary one about the same subjects. The two datasets may reflect similar but different latent grouping structures of the subjects. We propose an adaptive transfer clustering (ATC) algorithm that automatically leverages the commonality in the presence of unknown discrepancy, by optimizing an estimated bias-variance decomposition. It applies to a broad class of statistical models including Gaussian mixture models, stochastic block models, and latent class models. A theoretical analysis proves the optimality of ATC under the Gaussian mixture model and explicitly quantifies the benefit of transfer. Extensive simulations and real data experiments confirm our method's effectiveness in various scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。