arXiv:2502.11926cs.CL2025-02ACL被引 96

构建28种语言的多标签情感数据集,填补低资源语言研究空白。

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

论文配图:BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
图 1 · 摘自论文原文
  • 收集28种语言、多领域文本,由母语者标注情绪标签。
  • 跨语言与单语言实验验证数据集有效性,提升低资源语言性能。
  • 适合做多语言情感分析、跨语言迁移学习的研究者使用。

全球人们用复杂微妙的语言表达情感。尽管情感识别是自然语言处理中多项任务的统称,且对众多应用具有重要影响,但现有研究主要集中于高资源语言,导致低资源语言在研究投入和解决方案上存在显著差距,普遍缺乏高质量标注数据集。本文提出BRIGHTER——一个涵盖28种语言、多领域、多标签情感标注的数据集集合,重点覆盖非洲、亚洲、东欧及拉丁美洲的低资源语言,所有样本均由母语者标注。我们阐述了数据收集与标注过程中的挑战,并报告了单语言与跨语言多标签情感识别以及情感强度识别的实验结果。通过分析不同语言与文本领域下的性能差异(是否使用大模型),证明BRIGHTER数据集为解决文本情感识别中的语言鸿沟迈出了有意义的一步。

原文摘要 · Abstract (English)

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition--an umbrella term for several NLP tasks--impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities in research efforts and proposed solutions, particularly for under-resourced languages, which often lack high-quality annotated datasets. In this paper, we present BRIGHTER--a collection of multi-labeled, emotion-annotated datasets in 28 different languages and across several domains. BRIGHTER primarily covers low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. We highlight the challenges related to the data collection and annotation processes, and then report experimental results for monolingual and crosslingual multi-label emotion identification, as well as emotion intensity recognition. We analyse the variability in performance across languages and text domains, both with and without the use of LLMs, and show that the BRIGHTER datasets represent a meaningful step towards addressing the gap in text-based emotion recognition.

情感识别多语言低资源语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。