arXiv:2607.00143cs.CLcs.AI2026-07

构建多语种仇恨言论数据集并开发先进分析模型。

Hate Speech Detection in Turkish and Arabic: A Comprehensive Study

  • 构建土耳其五主题、阿拉伯一主题的仇恨言论数据集。
  • 提出基于BERT的模型,支持分类、强度预测等多维度分析。
  • 适合研究跨语言仇恨言论与内容安全的学者和工程师。

在线仇恨言论与全球针对少数群体的暴力事件上升相关,包括大规模枪击、私刑和种族清洗。在仇恨言论针对宗教、种族、民族、文化、国籍或移民身份时,社会面临言论自由与有效内容管理的平衡挑战。为此,我们构建了一个涵盖五个土耳其主题(难民、以巴冲突、反希腊情绪、亚列维派、亚美尼亚人、阿拉伯人、犹太人、库尔德人,以及LGBTI+)和一个阿拉伯主题(难民)的综合性仇恨言论数据集。同时,我们开发了基于BERT的先进模型,用于仇恨类别分类、仇恨强度预测、目标识别和仇恨言论片段检测,实现对网络言论中仇恨内容的全面分析。

原文摘要 · Abstract (English)

Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.

仇恨言论多语言BERT数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。