构建十年阿拉伯女性赋权语料库,分析社交媒体情感与社会关注。
Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus
- 收集2013-2024年77国5万余页的25万条阿拉伯语帖文。
- 涵盖超2.67亿次互动,含分享、评论与情绪反应数据。
- 支持阿拉伯语自然语言处理与社会情感研究,适合社科与计算学者。
本文发布阿拉伯女性与社会语料库(Arabic Women and Society Corpus),是2013至2024年间从77个国家51,660个页面收集的252,487条公开阿拉伯语Facebook帖子,覆盖超过2.67亿次用户互动。每条帖子包含分享、评论及情感反应等参与指标,揭示受众情绪与社会关注度。数据通过自动化流程完成语言识别、归一化与元数据清洗,确保可靠性与可复现性。该语料库支持大规模性别话语、社会改革与跨方言情感互动分析,服务于阿拉伯语自然语言处理、计算社会科学与数字传播研究。数据集及配套文档将按需开放用于学术研究。
原文摘要 · Abstract (English)
This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages across 77 countries between 2013 and 2024, resulting in more than 267 million user interactions. Each post includes engagement metrics such as shares, comments, and emotional reactions, providing a unique view of audience sentiment and social attention. The data were processed using an automated pipeline with language identification, normalization, and metadata cleaning to ensure reliability and reproducibility. The corpus enables large scale analysis of gender discourse, social reform, and emotional engagement across Arabic dialects. It supports research in Arabic natural language processing, computational social science, and digital communication studies. The dataset and accompanying documentation will be released under request for research use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。