arXiv:2502.17020cs.LGcs.AI2025-02被引 2

通过多分辨率分析,揭示短文本聚类中集群随数量变化的动态特性。

Moving Past Single Metrics: Exploring Short-Text Clustering Across Multiple Resolutions

  • 提出比例稳定性度量,评估不同聚类数下集群的持续性
  • 在3万条政治类推特简介上验证,集群随数量增加呈现可追踪的分化与重组
  • 用桑基图可视化集群演变,帮助理解数据内在结构

聚类数目通常在聚类问题中作为初始参数设定,虽具重要影响,但选择常难以解释。受生物信息学启发,本研究探讨聚类数目变化对集群性质的影响,提出一种评估聚类鲁棒性的方法,并系统化确定聚类数的策略。研究聚焦于短文本聚类,使用30,000条政治类推特简介数据,其中文本间词共现稀疏,难以发现有意义的聚类。引入比例稳定性度量,以揭示特定集群在不同聚类分辨率间的稳定性,并通过桑基图可视化结果,提供一种交互式工具以理解数据集特性。该可视化直观展示聚类数增加时集群的细分与重组过程,揭示了静态单一分辨率指标无法捕捉的深层信息。结果表明,不应追求单一‘最优’解,而应在信息量与复杂性之间权衡选择聚类数。

原文摘要 · Abstract (English)

Cluster number is typically a parameter selected at the outset in clustering problems, and while impactful, the choice can often be difficult to justify. Inspired by bioinformatics, this study examines how the nature of clusters varies with cluster number, presenting a method for determining cluster robustness, and providing a systematic method for deciding on the cluster number. The study focuses specifically on short-text clustering, involving 30,000 political Twitter bios, where the sparse co-occurrence of words between texts makes finding meaningful clusters challenging. A metric of proportional stability is introduced to uncover the stability of specific clusters between cluster resolutions, and the results are visualised using Sankey diagrams to provide an interrogative tool for understanding the nature of the dataset. The visualisation provides an intuitive way to track cluster subdivision and reorganisation as cluster number increases, offering insights that static, single-resolution metrics cannot capture. The results show that instead of seeking a single 'optimal' solution, choosing a cluster number involves balancing informativeness and complexity.

聚类分析短文本可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。