用表情符号和词语自动标注非英语情感,降低人工成本。
Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana
- 利用情感表情符号和词汇实现跨语言情感标注。
- 在英文、塞佩迪语、茨瓦纳语上准确率分别达66%、69%、63%。
- 仅需人工校对约三分之一的标签,适合资源匮乏语言。
情感分析有助于在人工智能助力社会公益、教育或营销等领域自动分析观点与情绪。尽管许多系统针对英文开发,但诸多非洲语言因缺乏数字语料(如带情感标签的文本)被归为低资源语言。主要原因是人工标注耗时费力。因此亟需自动且快速的方法以减少人工干预,提升标注效率。本文提出并分析一种基于情感表情符号与词汇信息的跨语言情感标注方法。实验基于SAfriSenti多语言语料库中英文、塞佩迪语和茨瓦纳语的推文。结果表明,该方法在英文推文上准确率达66%,塞佩迪语为69%,茨瓦纳语为63%,平均仅需人工修正约34%的标签。
原文摘要 · Abstract (English)
Sentiment analysis is a helpful task to automatically analyse opinions and emotions on various topics in areas such as AI for Social Good, AI in Education or marketing. While many of the sentiment analysis systems are developed for English, many African languages are classified as low-resource languages due to the lack of digital language resources like text labelled with corresponding sentiment classes. One reason for that is that manually labelling text data is time-consuming and expensive. Consequently, automatic and rapid processes are needed to reduce the manual effort as much as possible making the labelling process as efficient as possible. In this paper, we present and analyze an automatic language-independent sentiment labelling method that leverages information from sentiment-bearing emojis and words. Our experiments are conducted with tweets in the languages English, Sepedi and Setswana from SAfriSenti, a multilingual sentiment corpus for South African languages. We show that our sentiment labelling approach is able to label the English tweets with an accuracy of 66%, the Sepedi tweets with 69%, and the Setswana tweets with 63%, so that on average only 34% of the automatically generated labels remain to be corrected.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。