arXiv:2508.16046cs.IR2025-08被引 2

用LDA与聚类分析文章摘要,自动提取有效主题关键词

Estimating the Effective Topics of Articles and journals Abstract Using LDA And K-Means Clustering Algorithm

  • 结合LDA与K-Means算法挖掘文本主题,辅以WordNet提取关键词
  • 在关键词提取任务中,K-Means与LDA表现最优,准确率更高
  • 帮助研究者构建精准检索式,避免因主题误解导致的搜索偏差

随着文本文档数量激增,利用主题建模与文本聚类分析期刊和文章摘要已成为现代解决方案。主题建模与文本聚类相互促进,能有效管理海量文本。本研究采用LDA、K-Means聚类及词汇数据库WordNet进行关键词提取。实验表明,K-Means与LDA在关键词提取任务中表现最稳定可靠。该方法可辅助研究人员基于期刊与文章主题构建精准检索串,减少因主题理解偏差引发的检索误差。

原文摘要 · Abstract (English)

Analyzing journals and articles abstract text or documents using topic modelling and text clustering has become a modern solution for the increasing number of text documents. Topic modelling and text clustering are both intensely involved tasks that can benefit one another. Text clustering and topic modelling algorithms are used to maintain massive amounts of text documents. In this study, we have used LDA, K-Means cluster and also lexical database WordNet for keyphrases extraction in our text documents. K-Means cluster and LDA algorithms achieve the most reliable performance for keyphrase extraction in our text documents. This study will help the researcher to make a search string based on journals and articles by avoiding misunderstandings.

主题建模文本聚类关键词提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。