用MARBERT模型分析阿拉伯语推文情感与垃圾信息,提升客服响应质量
Spam and Sentiment Detection in Arabic Tweets Using MARBERT Model
- 基于MARBERT的深度学习模型处理阿拉伯语推文
- 在24,513条推文中实现高精度情感分类,负向样本占比超五成
- 适合关注中东地区社交媒体舆情分析的研究者与企业
沙特电信公司(STC)是沙特最受欢迎的企业之一,拥有大量客户。然而,用户满意度仍有提升空间。社交媒体是衡量用户满意度和情绪的重要平台,其中推特尤为突出。由于STC客服账号响应迅速,用户更倾向于通过推特表达反馈。为此,本研究采用情感分析技术,基于24,513条阿拉伯语推文数据集(含1,437条正面、13,828条负面、5,694条中性、1,221条讽刺、2,297条不确定推文),使用MARBERT模型进行训练,并以f1-score、精确率和召回率评估性能。该方法旨在识别用户情绪与垃圾信息,从而优化客户服务。实验结果表明,所提方案在准确性上优于现有技术。
原文摘要 · Abstract (English)
Saudi Telecom Company (STC) is among the most popular companies in Saudi Arabia, with many customers. Yet, there is still a big room for improvement in users' satisfaction. Social media is the most robust platform to gauge users' satisfaction and determine their sentiments and critics. Twitter is among the most popular social media platform in this regard. STC customers prefer to use Twitter to write their feedback because it's a fast way to get responses due to the STC customer services account. One way to achieve customer demands and improve customer service is using the Sentiment Analysis tool. Sentiment Analysis on Twitter is highly used because of the significant number of tweets and the different opinions. Likewise, Deep learning is the best existing Sentiment Analysis method, and it has diverse models. Bidirectional Encoder Representations from Transformers (BERT) model is one of the deep learning models which have achieved excellent results in Sentiment Analysis for Natural Language Processing (NLP). NLP is mainly investigated in the English language. However, for Arabic, there is a significant gap to be filled. This study trained the proposed model using MARBERT and measured the performance using f1-score, precision, and recall metrics. We trained the model with an Arabic dataset of 24,513 tweets, including 1,437 positive, 13,828 negative, 5,694 neutral, 1,221 sarcasm, and 2,297 indeterminate tweets. The main goal is to analyze the tweets and get the sentiment to improve STC customer service. The proposed scheme is promising in terms of accuracy in contrast to existing techniques in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。