arXiv:2502.16139cs.LGcs.CL2025-02被引 1

用微调词嵌入提升大规模文本聚类效果,性能显著优于传统方法。

An Improved Deep Learning Model for Word Embeddings Based Clustering for Large Text Datasets

  • 融合微调上下文嵌入与降维优化,改进聚类算法框架。
  • 聚类轮廓系数中位数提升45%(K-means)和67%(层次聚类)。
  • 适合需要高语义理解的海量文本分析任务。

本文提出一种基于微调词嵌入的改进型文本聚类技术,以WEClustering为基础模型,进一步引入微调的上下文嵌入、先进的降维方法及聚类算法优化。在基准数据集上的实验表明,该方法在轮廓系数、纯度和调整兰德指数(ARI)等指标上均有显著提升。其中,基于K均值的WE-Clustering_K++与基于层次聚类的WEClustering_A++,其轮廓系数中位数分别提升了45%和67%。该技术有助于弥合大规模文本挖掘中语义理解与统计稳健性之间的差距。

原文摘要 · Abstract (English)

In this paper, an improved clustering technique for large textual datasets by leveraging fine-tuned word embeddings is presented. WEClustering technique is used as the base model. WEClustering model is fur-ther improvements incorporating fine-tuning contextual embeddings, advanced dimensionality reduction methods, and optimization of clustering algorithms. Experimental results on benchmark datasets demon-strate significant improvements in clustering metrics such as silhouette score, purity, and adjusted rand index (ARI). An increase of 45% and 67% of median silhouette score is reported for the proposed WE-Clustering_K++ (based on K-means) and WEClustering_A++ (based on Agglomerative models), respec-tively. The proposed technique will help to bridge the gap between semantic understanding and statistical robustness for large-scale text-mining tasks.

文本聚类词嵌入深度学习降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。