arXiv:2605.22081cs.CL2026-05中稿 · LREC 2026 Main Con…

构建十年阿拉伯语社交媒体歧视数据集,支持多维度分析与公平性研究。

ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination

  • 收集293K条阿拉伯语社交帖子,融合点赞、评论等互动信号。
  • 涵盖200个关键词及20类歧视轴线,支持细粒度歧视识别。
  • 适合研究阿拉伯语公平性、平台生态与社会偏见的学者使用。

我们提出 ArabDiscrim,一个涵盖2014至2024年共293,000条公开阿拉伯语Facebook帖子的十年期词汇资源与语料库,聚焦种族主义与歧视议题。与现有以Twitter为主的语料库不同,ArabDiscrim整合了平台原生的互动信号(如点赞、分享、评论、页面元数据),支持语言与受众反应的联合分析。该资源包含200个经筛选的术语(100个与种族主义相关,100个与歧视相关),每个词根均具备13种以上形态变体,以及20个歧视轴线,用于捕捉基于身份的不平等待遇。同时提供明确的归因模式。数据在受限研究许可下发布,确保符合平台条款。ArabDiscrim可支持弱监督学习、轴向感知采样及平台生态研究,通过结合词汇深度与生态有效性,为面向公平性的、平台感知的阿拉伯语自然语言处理奠定基础。

原文摘要 · Abstract (English)

We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination. Unlike existing Twitter-centric datasets, ArabDiscrim integrates platform-native engagement signals, including reactions, shares, comments, and page metadata, enabling joint analysis of language and audience response. The resource includes 200 curated terms (100 racism-related and 100 discrimination-related) with morphological regex families (13+ inflections per lemma), and 20 discrimination axes capturing identity-based grounds for unequal treatment. It also provides explicit attribution patterns. Released under a restricted research-use license for ethical compliance with platform terms, ArabDiscrim supports weak supervision, axis-aware sampling, and platform ecology research. By bridging lexical depth and ecological validity, it establishes a foundation for fairness-oriented, platform-aware Arabic NLP.

社交媒体歧视检测阿拉伯语NLP数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。