首个面向社交媒体的细粒度情绪分析数据集,解决情绪类别复杂与数据稀缺难题。
EmoGRACE: Aspect-based emotion analysis for social media data
- 基于情感层级理论构建2621条推文标注数据集,支持五类情绪识别。
- 微调BERT模型在情绪分类任务上达到46.9%联合准确率,术语提取达70.1%F1。
- 适合从事情感计算、社交媒体分析的研究者参考使用。
尽管情感分析已从句子级别发展到方面级别(即识别具体的情感相关词),但方面级情绪分析(ABEA)仍面临数据集匮乏和情绪类别复杂性远高于二元情感的问题。本文首次构建了一个包含2,621条英文推文的ABEA训练数据集,并对基于BERT的模型进行微调,以完成方面术语抽取(ATE)和方面情绪分类(AEC)两个子任务。数据标注基于Shaver等人的层级情感理论,采用群体标注与多数投票策略以保证标签一致性。最终数据集包含愤怒、悲伤、快乐、恐惧及无情绪共五类方面级情绪标签。在此基础上,将最新ABS A模型GRACE用于ABEA任务,结果显示:ATE任务的F1分数为70.1%,联合提取任务(ATE+AEC)的准确率为46.9%。性能瓶颈主要归因于训练数据量小与任务复杂度高,导致模型过拟合,泛化能力有限。
原文摘要 · Abstract (English)
While sentiment analysis has advanced from sentence to aspect-level, i.e., the identification of concrete terms related to a sentiment, the equivalent field of Aspect-based Emotion Analysis (ABEA) is faced with dataset bottlenecks and the increased complexity of emotion classes in contrast to binary sentiments. This paper addresses these gaps, by generating a first ABEA training dataset, consisting of 2,621 English Tweets, and fine-tuning a BERT-based model for the ABEA sub-tasks of Aspect Term Extraction (ATE) and Aspect Emotion Classification (AEC). The dataset annotation process was based on the hierarchical emotion theory by Shaver et al. [1] and made use of group annotation and majority voting strategies to facilitate label consistency. The resulting dataset contained aspect-level emotion labels for Anger, Sadness, Happiness, Fear, and a None class. Using the new ABEA training dataset, the state-of-the-art ABSA model GRACE by Luo et al. [2] was fine-tuned for ABEA. The results reflected a performance plateau at an F1-score of 70.1% for ATE and 46.9% for joint ATE and AEC extraction. The limiting factors for model performance were broadly identified as the small training dataset size coupled with the increased task complexity, causing model overfitting and limited abilities to generalize well on new data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。