arXiv:2505.23297cs.CL2025-05EMNLP

首个乌克兰语情感识别基准数据集,填补了该语言情感分析空白。

EmoBench-UA: A Benchmark Dataset for Emotion Detection in Ukrainian

  • 基于英文情感标注规范构建乌克兰语标注体系
  • 通过众包平台获取高质量标注,覆盖多种文本类型
  • 验证了主流模型在非主流语言上的性能瓶颈,推动本地化研究

尽管乌克兰语自然语言处理在诸多文本任务上取得进展,但情感分类仍属研究空白,尚无公开可用的基准数据集。本文提出EmoBench-UA,首个乌克兰语文本情感检测的标注数据集。其标注方案参考了先前以英语为中心的情感识别工作(Mohammad et al., 2018; Mohammad, 2022)的指南。数据通过Toloka.ai平台进行众包标注,确保标注质量。我们在此数据集上评估了多种方法,包括基于语言学的基线模型、从英语翻译生成的合成数据,以及大型语言模型(LLMs)。结果表明,乌克兰语等非主流语言在情感分类上面临显著挑战,凸显了开发专用模型与训练资源的必要性。

原文摘要 · Abstract (English)

While Ukrainian NLP has seen progress in many texts processing tasks, emotion classification remains an underexplored area with no publicly available benchmark to date. In this work, we introduce EmoBench-UA, the first annotated dataset for emotion detection in Ukrainian texts. Our annotation schema is adapted from the previous English-centric works on emotion detection (Mohammad et al., 2018; Mohammad, 2022) guidelines. The dataset was created through crowdsourcing using the Toloka.ai platform ensuring high-quality of the annotation process. Then, we evaluate a range of approaches on the collected dataset, starting from linguistic-based baselines, synthetic data translated from English, to large language models (LLMs). Our findings highlight the challenges of emotion classification in non-mainstream languages like Ukrainian and emphasize the need for further development of Ukrainian-specific models and training resources.

情感识别乌克兰语数据集NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。