BERTopic在短篇印地语文本中表现优于传统主题模型。
BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
- 用上下文嵌入捕捉短文本语义关系,提升主题建模效果。
- 在多种主题数下,相干性得分均高于8种对比模型。
- 适合处理语言资源少的低资源短文本主题分析任务。
随着印地语等本土语言的短文本在现代媒体中日益增多,针对此类数据的鲁棒主题建模方法变得愈发重要。本研究探讨了BERTopic在印地语短文本主题建模中的表现,该领域在现有研究中尚属空白。利用上下文嵌入,BERTopic能够捕捉数据中的语义关系,相较于传统模型在短而多样的文本上更具优势。我们使用6种不同的文档嵌入模型评估BERTopic,并与8种成熟主题模型进行对比:潜在狄利克雷分配(LDA)、非负矩阵分解(NMF)、潜在语义索引(LSI)、主题模型加性正则化(ARTM)、概率潜在语义分析(PLSA)、嵌入主题模型(ETM)、联合主题模型(CTM)和Top2Vec。通过不同主题数量下的相干性得分进行评估。结果表明,BERTopic在生成连贯主题方面始终优于其他模型。
原文摘要 · Abstract (English)
As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modeling Hindi short texts, an area that has been under-explored in existing research. Using contextual embeddings, BERTopic can capture semantic relationships in data, making it potentially more effective than traditional models, especially for short and diverse texts. We evaluate BERTopic using 6 different document embedding models and compare its performance against 8 established topic modeling techniques, such as Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), Latent Semantic Indexing (LSI), Additive Regularization of Topic Models (ARTM), Probabilistic Latent Semantic Analysis (PLSA), Embedded Topic Model (ETM), Combined Topic Model (CTM), and Top2Vec. The models are assessed using coherence scores across a range of topic counts. Our results reveal that BERTopic consistently outperforms other models in capturing coherent topics from short Hindi texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。