构建北非多语种假新闻数据集,揭示情感化语言如何助推谣言传播。
BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon

- 整合真假新闻与评论,覆盖阿拉伯语、方言、法语等多语言
- 发现假新闻依赖情绪化叙事和混合语言提升传播力
- 适合研究虚假信息、低资源语言处理与社会舆情的学者
社交媒体上假新闻的快速传播已成为重大挑战,尤其在北非这类多语种且资源匮乏的地区。本文提出BOUTEF,一个大规模多语种语料库,用于研究阿尔及利亚和突尼斯的假新闻传播特征与影响。该语料库包含真假叙事、用户评论及经验证的辟谣信息,覆盖标准阿拉伯语、阿尔及利亚/突尼斯方言、Arabizi、法语、英语及混用语言。基于此,我们采用量化与质性结合的方法,分析主题分布、语言策略、情感模式与社交互动。统计结果显示,主题类别与内容真伪显著相关,用户参与度与假新闻可见性高度正相关。假新闻普遍使用情绪化叙事、煽动性框架与混合语言实践以增强传播力,而辟谣内容则更侧重事实性与验证。阿尔及利亚与突尼斯的对比分析揭示了共性与受社会政治背景影响的差异。研究强调非正式语言在信息失序中的作用。本工作提供了一个丰富、标注完整且公开可获取的数据集,推动假新闻检测、低资源语言处理及复杂语境下信息紊乱的研究。
原文摘要 · Abstract (English)
The rapid spread of fake news on social media has become a major challenge, particularly in multilingual and under-resourced contexts such as North Africa. In this paper, we introduce BOUTEF, a large-scale multilingual corpus designed to study the propagation, characteristics, and impact of fake news in Algeria and Tunisia. The corpus integrates three complementary components: fake narratives, genuine narratives, and associated user-generated comments, along with verified debunking information. It covers a wide range of languages and linguistic varieties, including MSA, Algerian and Tunisian dialects, Arabizi, French, English, and code-switched language. Building on this resource, we conduct a comprehensive empirical analysis combining quantitative and qualitative approaches. We examine thematic distributions, linguistic and rhetorical strategies, sentiment patterns, and social engagement dynamics. Statistical analyses reveal significant associations between thematic categories and message veracity, as well as strong correlations between user engagement and the visibility of fake content. Our findings show that fake news relies heavily on emotionally charged narratives, sensational framing, and hybrid linguistic practices that enhance virality and audience engagement. In contrast, debunking content adopts a more factual and verification-oriented style. Furthermore, a comparative analysis between Algeria and Tunisia highlights both shared dynamics and country-specific characteristics shaped by sociopolitical contexts. The results emphasize the role of informal language practices in the diffusion and reception of misinformation. By providing a rich, annotated, and publicly available dataset, this work contributes to advancing research on fake news detection, low-resource language processing, and the understanding of information disorders in complex linguistic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。