arXiv:2501.08284cs.CL2025-01NAACL被引 20

构建15种非洲语言的仇恨言论数据集,由本地人标注,提升跨文化语境下的识别准确性。

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

  • 基于15种非洲语言,由母语者按本地文化标注,确保语义理解准确。
  • 提供分类基线结果,验证大模型在多语言仇恨言论识别中的有效性。
  • 适合研究多语言社会计算、数字公平与本土化内容治理的学者和开发者。

仇恨言论和侮辱性语言是全球性现象,需结合社会文化背景才能理解、识别与管理。然而,在全球南方许多地区,因过度依赖脱离语境的关键词检测,导致(1)监管缺失与(2)误删误封并存。同时,高层人物常成为监管焦点,而针对少数群体的大规模仇恨攻击却常被忽视。这些问题主要源于本地语言高质量数据稀缺,以及缺乏本地社区参与数据收集、标注与治理。为此,我们推出AfriHate:一个涵盖15种非洲语言的仇恨言论与侮辱性语言数据集。所有样本均由熟悉当地文化的母语者标注。我们报告了数据构建中的挑战,并提供了使用与不使用大语言模型的分类基线结果。数据集、标注详情及仇恨言论词汇表已开源于https://github.com/AfriHate/AfriHate。

原文摘要 · Abstract (English)

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge to be understood, identified, and moderated. However, in many regions of the Global South, there have been several documented occurrences of (1) absence of moderation and (2) censorship due to the reliance on keyword spotting out of context. Further, high-profile individuals have frequently been at the center of the moderation process, while large and targeted hate speech campaigns against minorities have been overlooked. These limitations are mainly due to the lack of high-quality data in the local languages and the failure to include local communities in the collection, annotation, and moderation processes. To address this issue, we present AfriHate: a multilingual collection of hate speech and abusive language datasets in 15 African languages. Each instance in AfriHate is annotated by native speakers familiar with the local culture. We report the challenges related to the construction of the datasets and present various classification baseline results with and without using LLMs. The datasets, individual annotations, and hate speech and offensive language lexicons are available on https://github.com/AfriHate/AfriHate

仇恨言论多语言非洲语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。