构建中英文情感与希望话语双语数据集,助力低资源语言情感分析
EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic
- 构建含23,456条阿拉伯语和10,036条英语的双语标注数据集
- 标注涵盖情感强度、复杂性及原因,希望话语分类细致
- 适用于跨语言情感分析与低资源语言NLP研究
本研究提出一个双语数据集,包含23,456条阿拉伯语和10,036条英语样本,用于情感与希望话语标注,解决多情感(情绪与希望)数据稀缺问题。数据集提供情感强度、复杂性及成因的全面标注,并对希望话语进行详细分类与子类划分。为确保标注可靠性,采用Fleiss' Kappa评估,阿拉伯语与英语标注者间一致性达0.75-0.85。基准模型(机器学习)获得micro-F1-Score=0.67,验证了标注质量。该数据集为提升低资源语言自然语言处理能力提供了宝贵资源,推动跨语言情感与希望话语分析发展。
原文摘要 · Abstract (English)
This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datasets. The dataset provides comprehensive annotations capturing emotion intensity, complexity, and causes, alongside detailed classifications and subcategories for hope speech. To ensure annotation reliability, Fleiss' Kappa was employed, revealing 0.75-0.85 agreement among annotators both for Arabic and English language. The evaluation metrics (micro-F1-Score=0.67) obtained from the baseline model (i.e., using a machine learning model) validate that the data annotations are worthy. This dataset offers a valuable resource for advancing natural language processing in underrepresented languages, fostering better cross-linguistic analysis of emotions and hope speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。