构建首个面向波斯语推特的情感分析数据集,支持三类情感识别。
Exa-PSD: a new Persian sentiment analysis dataset on Twitter
- 从1.2万条波斯语推特中收集并标注数据,使用5名母语者保证质量。
- 基于ParsBERT和RoBERTa模型评估,取得79.87%的宏平均F1分数。
- 专为社交媒体中的口语化、反讽表达设计,适合波斯语情感分析研究者。
如今,推特等社交平台是人们交流的主要渠道,其中蕴含大量可挖掘的舆论信息。情感分析在自然语言处理中至关重要,用于识别个体对特定话题的观点。尽管预训练语言模型发展迅速,但波斯语自然语言处理仍面临诸多挑战。现有波斯语数据集多集中于产品、食品、酒店等特定领域,而社交媒体用户常使用反讽、方言等表达方式。为此,本文提出Exa-PSD数据集,从波斯语推特中收集12,000条已标注文本,由5位母语者进行标注,分为正面、中性、负面三类。我们分析了该数据集的统计特征,并采用预训练的ParsBERT和RoBERTa作为基线模型进行评估,获得79.87%的宏平均F1分数,表明该数据集与模型具有较高应用价值,可为波斯语情感分析系统提供有力支持。
原文摘要 · Abstract (English)
Today, Social networks such as Twitter are the most widely used platforms for communication of people. Analyzing this data has useful information to recognize the opinion of people in tweets. Sentiment analysis plays a vital role in NLP, which identifies the opinion of the individuals about a specific topic. Natural language processing in Persian has many challenges despite the adventure of strong language models. The datasets available in Persian are generally in special topics such as products, foods, hotels, etc while users may use ironies, colloquial phrases in social media To overcome these challenges, there is a necessity for having a dataset of Persian sentiment analysis on Twitter. In this paper, we introduce the Exa sentiment analysis Persian dataset, which is collected from Persian tweets. This dataset contains 12,000 tweets, annotated by 5 native Persian taggers. The aforementioned data is labeled in 3 classes: positive, neutral and negative. We present the characteristics and statistics of this dataset and use the pre-trained Pars Bert and Roberta as the base model to evaluate this dataset. Our evaluation reached a 79.87 Macro F-score, which shows the model and data can be adequately valuable for a sentiment analysis system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。