arXiv:2409.12840cs.CL2024-09被引 11

用词典法分析推文情感极性,比较多种模型表现

Lexicon-Based Sentiment Analysis on Text Polarities with Evaluation of Classification Models

  • 基于词典法在词级别识别情感强度和主观性
  • 随机森林准确率达81%为最佳模型
  • 适合对社交媒体情感分析感兴趣的实践者

情感分析在数字平台具有广泛应用潜力,可提取文本中的情感极性以理解其强度与主观性。本文采用词典法进行情感分析,并评估了在文本数据上训练的分类模型性能。词典法在词级别识别情感强度与主观性,通过分类确定文本中有信息量的词汇,并给出词汇极性的量化排名。研究基于多类问题,将文本标记为正面、负面或中性。使用包含160万条未处理推文的Twitter情感数据集,结合TextBlob和Vader Sentiment等词典法引入文本中性度衡量。词典分析揭示了词数与情感强度对文本分类的影响。对朴素贝叶斯、支持向量机、多项式逻辑回归、随机森林及极端梯度提升(XGBoost)模型进行了多指标对比分析,结果显示随机森林表现最优,准确率为81%。此外,还将情感分析应用于基于在线活动对推特账号进行人格判断的案例。

原文摘要 · Abstract (English)

Sentiment analysis possesses the potential of diverse applicability on digital platforms. Sentiment analysis extracts the polarity to understand the intensity and subjectivity in the text. This work uses a lexicon-based method to perform sentiment analysis and shows an evaluation of classification models trained over textual data. The lexicon-based methods identify the intensity of emotion and subjectivity at word levels. The categorization identifies the informative words inside a text and specifies the quantitative ranking of the polarity of words. This work is based on a multi-class problem of text being labeled as positive, negative, or neutral. Twitter sentiment dataset containing 1.6 million unprocessed tweets is used with lexicon-based methods like Text Blob and Vader Sentiment to introduce the neutrality measure on text. The analysis of lexicons shows how the word count and the intensity classify the text. A comparative analysis of machine learning models, Naiive Bayes, Support Vector Machines, Multinomial Logistic Regression, Random Forest, and Extreme Gradient (XG) Boost performed across multiple performance metrics. The best estimations are achieved through Random Forest with an accuracy score of 81%. Additionally, sentiment analysis is applied for a personality judgment case against a Twitter profile based on online activity.

情感分析词典法机器学习推特数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。