arXiv:2503.18253cs.CL2025-03被引 3

为埃塞俄比亚语言情感分析添加强度标注,提升多标签情绪识别效果。

Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages

  • 在EthioEmo数据集上添加情绪强度标注,支持多情绪强度判别。
  • 非洲本地化小模型在情绪分类任务中优于开源大模型。
  • 引入情绪强度特征可显著提升多标签情绪识别性能。

开发和集成情感理解模型对人机交互任务至关重要,涵盖客户反馈分析、市场研究及社交媒体监控。由于用户常在同一文本中表达多种情绪,采用多标签格式标注情感数据集对捕捉这种复杂性尤为关键。现有埃塞俄比亚语言的多语言、多标签情感数据集EthioEmo缺乏情绪强度标注,而该标注对区分情绪强弱程度极为重要。本文通过添加情绪强度标注扩展EthioEmo数据集,并在此基础上对当前主流的仅编码器预训练语言模型(PLMs)和大型语言模型(LLMs)进行基准测试。结果表明,面向非洲语境的编码器模型表现持续优于开源大模型,凸显了文化与语言适配的小模型在情感理解中的价值。同时,引入情绪强度特征能有效提升多标签情感分类性能。数据已公开于https://huggingface.co/datasets/Tadesse/EthioEmo-intensities。

原文摘要 · Abstract (English)

Developing and integrating emotion-understanding models are essential for a wide range of human-computer interaction tasks, including customer feedback analysis, marketing research, and social media monitoring. Given that users often express multiple emotions simultaneously within a single instance, annotating emotion datasets in a multi-label format is critical for capturing this complexity. The EthioEmo dataset, a multilingual and multi-label emotion dataset for Ethiopian languages, lacks emotion intensity annotations, which are crucial for distinguishing varying degrees of emotion, as not all emotions are expressed with the same intensity. We extend the EthioEmo dataset to address this gap by adding emotion intensity annotations. Furthermore, we benchmark state-of-the-art encoder-only Pretrained Language Models (PLMs) and Large Language Models (LLMs) on this enriched dataset. Our results demonstrate that African-centric encoder-only models consistently outperform open-source LLMs, highlighting the importance of culturally and linguistically tailored small models in emotion understanding. Incorporating an emotion-intensity feature for multi-label emotion classification yields better performance. The data is available at https://huggingface.co/datasets/Tadesse/EthioEmo-intensities.

情感分析多标签非洲语言强度标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。