arXiv:2502.08319cs.CL2025-02被引 3

首个阿拉伯语多标签宣传、情感与情绪数据集,助力中文读者理解中东舆论机制。

MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection

  • 构建包含8000篇新闻的多标签数据集,覆盖宣传、情感与情绪三维度。
  • 在该数据集上,GPT-4o-mini等大模型实现传播性识别准确率超85%。
  • 开源标注指南与代码,适合研究阿拉伯语舆情与多任务语言模型者使用。

宣传是一种自古以来用于通过修辞和心理策略影响公众意见的说服形式。尽管阿拉伯语是互联网上第四大使用语言,但针对英语以外语言的宣传检测资源,尤其是阿拉伯语,仍极度匮乏。为填补这一空白,本文提出首个多标签阿拉伯语宣传、情感与情绪检测数据集 MultiProSE,作为现有阿拉伯语宣传数据集 ArPro 的开源扩展,新增每条文本的情感与情绪标注。该数据集共包含8,000篇经人工标注的新闻文章,是目前最大的阿拉伯语宣传数据集。针对各任务,采用 GPT-4o-mini 等大语言模型(LLMs)及三种基于 BERT 的预训练语言模型(PLMs)构建多个基线模型。数据集、标注指南与源代码均已公开发布,以推动阿拉伯语语言模型的研究发展,并增进对新闻媒体中多重意见维度交互机制的理解。

原文摘要 · Abstract (English)

Propaganda is a form of persuasion that has been used throughout history with the intention goal of influencing people's opinions through rhetorical and psychological persuasion techniques for determined ends. Although Arabic ranked as the fourth most-used language on the internet, resources for propaganda detection in languages other than English, especially Arabic, remain extremely limited. To address this gap, the first Arabic dataset for Multi-label Propaganda, Sentiment, and Emotion (MultiProSE) has been introduced. MultiProSE is an open-source extension of the existing Arabic propaganda dataset, ArPro, with the addition of sentiment and emotion annotations for each text. This dataset comprises 8,000 annotated news articles, which is the largest propaganda dataset to date. For each task, several baselines have been developed using large language models (LLMs), such as GPT-4o-mini, and pre-trained language models (PLMs), including three BERT-based models. The dataset, annotation guidelines, and source code are all publicly released to facilitate future research and development in Arabic language models and contribute to a deeper understanding of how various opinion dimensions interact in news media1.

阿拉伯语多标签舆情分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。