回顾2013-2021年SemEval竞赛的顶级情感分析系统,梳理技术演进路径。
Sentiment Analysis in SemEval: A Review of Sentiment Identification Approaches
- 分析658支队伍在6届SemEval中的高分系统,追踪方法演进
- 从词典法到词嵌入,神经网络与变压器主导分类阶段
- 为研究者提供可复用方案,助力新系统快速构建
社交媒体平台已成为社交互动和意见表达的重要载体。情感分析技术致力于实现对生成数据中情感、情绪及讨论主题的检索与分析。国际语义评估研讨会(SemEval)吸引了众多研究者与从业者参与,聚焦于构建情感分析系统。本文研究了2013至2021年间各届SemEval的顶尖系统,期间共有658支团队参与,关注度逐年上升。我们分析了所提系统的共性,揭示了情感分析系统核心组件——数据获取、预处理与分类——的发展趋势。结果显示,预处理技术被广泛采用,特征工程与词表示方式由词典基础转向词嵌入,分类阶段则以神经网络和变压器模型为主导,推动了现成模型的广泛应用。研究还为研究人员提供了基于实证系统的洞察,有助于新系统的快速原型开发,并支持从业者为未来SemEval竞赛做准备。
原文摘要 · Abstract (English)
Social media platforms are becoming the foundations of social interactions including messaging and opinion expression. In this regard, Sentiment Analysis techniques focus on providing solutions to ensure the retrieval and analysis of generated data including sentiments, emotions, and discussed topics. International competitions such as the International Workshop on Semantic Evaluation (SemEval) have attracted many researchers and practitioners with a special research interest in building sentiment analysis systems. In our work, we study top-ranking systems for each SemEval edition during the 2013-2021 period, a total of 658 teams participated in these editions with increasing interest over years. We analyze the proposed systems marking the evolution of research trends with a focus on the main components of sentiment analysis systems including data acquisition, preprocessing, and classification. Our study shows an active use of preprocessing techniques, an evolution of features engineering and word representation from lexicon-based approaches to word embeddings, and the dominance of neural networks and transformers over the classification phase fostering the use of ready-to-use models. Moreover, we provide researchers with insights based on experimented systems which will allow rapid prototyping of new systems and help practitioners build for future SemEval editions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。