对比多种模型,发现本地化BERT能更好挖掘马拉地语主题
Topic Modeling in Marathi
- 用BERTopic框架结合印地语专用BERT模型进行主题建模
- 本地化BERT+BERTopic在主题连贯性和多样性上优于LDA
- 为马拉地语等印地语系语言提供可复现的主题分析方案
尽管英语主题建模已广泛研究,但印地语系语言的主题建模仍属罕见。资源匮乏、语言结构多样及独特挑战导致相关研究稀缺。本文针对马拉地语,比较了多种主题建模方法,包括多语言与单语言BERT模型,以及非BERT方法,以主题连贯性与主题多样性为评估指标。结果表明,使用印地语训练的BERT模型与BERTopic结合,显著优于传统LDA模型。该方法为马拉地语等印地语系语言的主题分析提供了有效路径。
原文摘要 · Abstract (English)
While topic modeling in English has become a prevalent and well-explored area, venturing into topic modeling for Indic languages remains relatively rare. The limited availability of resources, diverse linguistic structures, and unique challenges posed by Indic languages contribute to the scarcity of research and applications in this domain. Despite the growing interest in natural language processing and machine learning, there exists a noticeable gap in the comprehensive exploration of topic modeling methodologies tailored specifically for languages such as Hindi, Marathi, Tamil, and others. In this paper, we examine several topic modeling approaches applied to the Marathi language. Specifically, we compare various BERT and non-BERT approaches, including multilingual and monolingual BERT models, using topic coherence and topic diversity as evaluation metrics. Our analysis provides insights into the performance of these approaches for Marathi language topic modeling. The key finding of the paper is that BERTopic, when combined with BERT models trained on Indic languages, outperforms LDA in terms of topic modeling performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。