arXiv:2504.04314cs.CLcs.AI2025-04被引 2

用大模型找最佳聚类数,让结果既准确又易懂。

Balancing Complexity and Informativeness in LLM-Based Clustering: Finding the Goldilocks Zone

  • 用LLM生成聚类名称,从语义密度和信息论评估效果。
  • 16-22个聚类时平衡了准确性和可解释性,表现最佳。
  • 发现聚类命名与内容相似度直接影响分类准确率。

短文本聚类面临信息量与可解释性之间的权衡难题。传统评估指标常忽略此矛盾。受语言学沟通效率启发,本文通过量化信息量与认知简洁性的权衡,探究最优聚类数量。利用大语言模型(LLM)生成聚类名称,并基于语义密度、信息论和聚类准确性进行评估。结果表明,在LLM生成的嵌入上使用高斯混合模型(GMM)聚类,相比随机分配显著提升语义密度,有效归类相似个人简介。然而随着聚类数增加,可解释性下降,表现为生成式LLM根据聚类名称正确分配简介的能力减弱。逻辑回归分析证实,分类准确率取决于简介与其所属聚类名称的语义相似度,以及与其他聚类的区别程度。研究揭示出一个‘恰到好处’的区间:16至22个聚类,与语言学中的词汇分类效率相呼应。该发现为理论建模与实际应用提供指导,推动未来研究优化聚类的可解释性与实用性。

原文摘要 · Abstract (English)

The challenge of clustering short text data lies in balancing informativeness with interpretability. Traditional evaluation metrics often overlook this trade-off. Inspired by linguistic principles of communicative efficiency, this paper investigates the optimal number of clusters by quantifying the trade-off between informativeness and cognitive simplicity. We use large language models (LLMs) to generate cluster names and evaluate their effectiveness through semantic density, information theory, and clustering accuracy. Our results show that Gaussian Mixture Model (GMM) clustering on embeddings generated by a LLM, increases semantic density compared to random assignment, effectively grouping similar bios. However, as clusters increase, interpretability declines, as measured by a generative LLM's ability to correctly assign bios based on cluster names. A logistic regression analysis confirms that classification accuracy depends on the semantic similarity between bios and their assigned cluster names, as well as their distinction from alternatives. These findings reveal a "Goldilocks zone" where clusters remain distinct yet interpretable. We identify an optimal range of 16-22 clusters, paralleling linguistic efficiency in lexical categorization. These insights inform both theoretical models and practical applications, guiding future research toward optimising cluster interpretability and usefulness.

聚类优化大模型应用可解释性语义密度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。