arXiv:2412.03584cs.SIcs.LG2024-12被引 1

提出新聚类相似性度量ResMI,更鲁棒且可解释。

Resampled Mutual Information for Clustering and Community Detection

  • 融合信息论与成对计数思想,不需调整项
  • 在高簇数和非对称分布下表现更稳健
  • 适合真实社交网络社区发现任务

我们提出一种新的聚类相似性度量——重采样互信息(Resampled Mutual Information, ResMI),该方法结合了信息论与成对计数方法的优点。与基于随机修正的度量类似,ResMI满足常数基线性质,但无需调整项,且完全可在信息论语言中解释。在合成数据集上的实验表明,ResMI对现有度量常见的偏差具有鲁棒性,尤其在高簇数和非对称簇分布场景下表现优异。此外,我们在两个真实的接触追踪网络上验证了ResMI能有效识别有意义的社区结构。

原文摘要 · Abstract (English)

We introduce resampled mutual information (ResMI), a novel measure of clustering similarity that combines insights from information theoretic and pair counting approaches to clustering and community detection. Similar to chance-corrected measures, ResMI satisfies the constant baseline property, but it has the advantages of not requiring adjustment terms and being fully interpretable in the language of information theory. Experiments on synthetic datasets demonstrate that ResMI is robust to common biases exhibited by existing measures, particularly in settings with high cluster counts and asymmetric cluster distributions. Additionally, we show that ResMI identifies meaningful community structures in two real contact tracing networks.

聚类评估社区检测信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。