arXiv:2510.09799cs.DScs.DC2025-10

解决多个机构在部分重叠特征上分布式聚类的问题

Distributed clustering in partially overlapping feature spaces

  • 采用联邦与单次通信两种算法,支持本地灵活聚类
  • 在三个公开数据集上验证性能,收敛至中心化最优解
  • 适合医疗等多机构协作场景,保护数据隐私

我们提出并解决了一种新型分布式聚类问题:每个参与方拥有仅包含全部特征子集的私有数据集,且部分特征在多个数据集中出现。该场景广泛存在于医疗等领域,不同机构对相似患者提供互补数据。本文提出两种算法应对此类特征空间异构问题:第一种为联邦算法,各参与方协同更新全局聚类中心;第二种为单次通信算法,各参与方将本地聚类的统计参数发送给中央服务器,由其生成并合并合成代理数据集。两种方法均允许参与者使用任意本地聚类算法,实现灵活性与个性化计算成本。通过假设本地数据源自对初始集中数据的分割与掩码,我们识别出算法收敛至最优中心化解的条件。最后,在三个公开数据集上测试了算法的实际性能。

原文摘要 · Abstract (English)

We introduce and address a novel distributed clustering problem where each participant has a private dataset containing only a subset of all available features, and some features are included in multiple datasets. This scenario occurs in many real-world applications, such as in healthcare, where different institutions have complementary data on similar patients. We propose two different algorithms suitable for solving distributed clustering problems that exhibit this type of feature space heterogeneity. The first is a federated algorithm in which participants collaboratively update a set of global centroids. The second is a one-shot algorithm in which participants share a statistical parametrization of their local clusters with the central server, who generates and merges synthetic proxy datasets. In both cases, participants perform local clustering using algorithms of their choice, which provides flexibility and personalized computational costs. Pretending that local datasets result from splitting and masking an initial centralized dataset, we identify some conditions under which the proposed algorithms are expected to converge to the optimal centralized solution. Finally, we test the practical performance of the algorithms on three public datasets.

分布式聚类联邦学习医疗数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。