构建埃塞俄比亚四语多标签情感数据集,评估大模型跨语言情感理解能力
Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding
- 构建EthioEmo数据集,覆盖阿姆哈拉语、奥罗莫语等四种埃塞俄比亚语言
- 高资源英语仍难实现精准多标签情感分类,低资源语言性能更差
- 验证不同模型架构在多语言情感任务中的差异,适合跨语言情感研究者
大型语言模型(LLMs)展现出强大的学习与推理能力。然而,相较于其他自然语言处理任务,多语言、多标签情感评估在大模型中仍研究不足。本文提出EthioEmo,一个包含四种埃塞俄比亚语言(阿姆哈拉语amh、奥罗莫语orm、索马里语som、提格利尼亚语tir)的多标签情感分类数据集。我们在额外的SemEval 2018 Task 1英文多标签情感数据集上进行广泛实验,涵盖编码器-编码器、编码器-解码器和解码器-only三种语言模型结构。对比了零样本与小样本方法与微调小型模型的效果。结果表明,即使在高资源语言如英语中,准确的多标签情感分类依然不足,且高资源与低资源语言间存在显著性能差距。不同语言与模型类型的表现也各异。EthioEmo已公开,以促进对语言模型情感理解及人类跨语言情感表达机制的研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages, namely, Amharic (amh), Afan Oromo (orm), Somali (som), and Tigrinya (tir). We perform extensive experiments with an additional English multi-label emotion dataset from SemEval 2018 Task 1. Our evaluation includes encoder-only, encoder-decoder, and decoder-only language models. We compare zero and few-shot approaches of LLMs to fine-tuning smaller language models. The results show that accurate multi-label emotion classification is still insufficient even for high-resource languages such as English, and there is a large gap between the performance of high-resource and low-resource languages. The results also show varying performance levels depending on the language and model type. EthioEmo is available publicly to further improve the understanding of emotions in language models and how people convey emotions through various languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。