为低资源库尔德语构建BERT模型,提升情感分析性能
KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
- 基于BERT架构构建中央库尔德语专用语言模型
- 在情感分析任务上超越传统Word2Vec方法
- 为低资源语言自然语言处理提供新范式
本文通过将双向编码器表示(BERT)融入自然语言处理技术,提升对中央库尔德语的情感分析研究。库尔德语属于低资源语言,语言多样性高但计算资源有限,情感分析难度较大。此前多采用如Word2Vec等传统词嵌入模型,但随着BERT等新型语言模型的出现,性能提升成为可能。BERT更优的词向量表示能力有助于捕捉库尔德语中细微的语义信息与上下文特征,从而为低资源语言的情感分析设立新基准。
原文摘要 · Abstract (English)
This paper enhances the study of sentiment analysis for the Central Kurdish language by integrating the Bidirectional Encoder Representations from Transformers (BERT) into Natural Language Processing techniques. Kurdish is a low-resourced language, having a high level of linguistic diversity with minimal computational resources, making sentiment analysis somewhat challenging. Earlier, this was done using a traditional word embedding model, such as Word2Vec, but with the emergence of new language models, specifically BERT, there is hope for improvements. The better word embedding capabilities of BERT lend to this study, aiding in the capturing of the nuanced semantic pool and the contextual intricacies of the language under study, the Kurdish language, thus setting a new benchmark for sentiment analysis in low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。