新框架可同时识别阿拉伯语方言与情绪,准确率超现有方法。
A Novel Dialect-Aware Framework for the Classification of Arabic Dialects and Emotions
- 分三模块构建方言与情绪联合识别框架
- 方言分类准确率达88.9%,情绪检测在埃及/海湾方言中达89.1%/79%
- 能生成方言感知的情感词典,适合多语言情感分析应用
阿拉伯语是现存最古老的语言之一,不同地区发展出独特方言。方言与情绪识别在阿拉伯文本分析中有重要应用,如根据用户评论判断其来源地,或让智能聊天机器人根据情绪响应。现有研究缺乏对不同方言中情绪表达方式的考量,本研究提出一种新型框架,可从文本中同时识别阿拉伯语方言与情绪。框架包含三个模块:文本预处理、分类和聚类,其中聚类模块具备构建新型方言感知情绪词典的能力。该框架生成了针对不同方言的新情感词典,在方言分类上达到88.9%准确率,比当前最优结果高出6.45个百分点;在埃及与海湾方言中,情绪检测准确率分别为89.1%和79%。
原文摘要 · Abstract (English)
Arabic is one of the oldest languages still in use today. As a result, several Arabic-speaking regions have developed dialects that are unique to them. Dialect and emotion recognition have various uses in Arabic text analysis, such as determining an online customer's origin based on their comments. Furthermore, intelligent chatbots that are aware of a user's emotions can respond appropriately to the user. Current research in emotion detection in the Arabic language lacks awareness of how emotions are exhibited in different dialects, which motivates the work found in this study. This research addresses the problems of dialect and emotion classification in Arabic. Specifically, this is achieved by building a novel framework that can identify and predict Arabic dialects and emotions from a given text. The framework consists of three modules: A text-preprocessing module, a classification module, and a clustering module with the novel capability of building new dialect-aware emotion lexicons. The proposed framework generated a new emotional lexicon for different dialects. It achieved an accuracy of 88.9% in classifying Arabic dialects, which outperforms the state-of-the-art results by 6.45 percentage points. Furthermore, the framework achieved 89.1-79% accuracy in detecting emotions in the Egyptian and Gulf dialects, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。