用大模型识别社交媒体中的隐晦反穆斯林言论
Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech
- 用LLM分析反穆斯林黑话在极端平台的语义
- 发现反穆斯林言论毒性评分高于反犹言论
- 揭示其与极右翼、阴谋论及移民议题的关联
近年来,数字通信网络推动了西方社会伊斯兰恐惧症的蔓延。本文大规模分析了4Chan、Gab、Telegram等极端社交平台上流传的半编码反穆斯林术语(如muzrat、pislam、mudslime、mohammedan、muzzies)。这些词汇在一般语境中看似中性,难以被人工或自动化系统准确识别为仇恨言论。研究首先利用大语言模型(LLMs)验证其对这些词汇的理解能力;其次,谷歌Perspective API显示,反穆斯林内容的毒性评分高于反犹主义等其他仇恨言论类别;最后采用BERT主题建模提取不同话题。结果表明,尽管LLMs能理解这些词形外词汇(OOV),但现有内容审核和算法检测仍需改进。主题分析进一步显示,反穆斯林言论广泛存在于政治、阴谋论及极右翼运动中,尤其针对穆斯林移民。本研究是首个系统分析半编码反穆斯林言论的工作,揭示了其全球传播特征。
原文摘要 · Abstract (English)
In recent years, Islamophobia has gained significant traction across Western societies, fueled by the rise of digital communication networks. This paper performs a large-scale analysis of specialized, semi-coded Islamophobic terms such as (muzrat, pislam, mudslime, mohammedan, muzzies) floated on extremist social platforms, i.e., 4Chan, Gab, Telegram, etc. Many of these terms appear lexically neutral or ambiguous outside of specific contexts, making them difficult for both human moderators and automated systems to reliably identify as hate speech. First, we use Large Language Models (LLMs) to show their ability to understand these terms. Second, Google Perspective API suggests that Islamophobic posts tend to receive higher toxicity scores than other categories of hate speech like Antisemitism. Finally, we use BERT topic modeling approach to extract different topics and Islamophobic discourse on these social platforms. Our findings indicate that LLMs understand these Out-Of-Vocabulary (OOV) slurs; however, further improvements in moderation strategies and algorithmic detection are necessary to address such discourse effectively. Our topic modeling also indicates that Islamophobic text is found across various political, conspiratorial, and far-right movements and is particularly directed against Muslim immigrants. Taken altogether, we performed one of the first studies on Islamophobic semi-coded terms and shed a global light on Islamophobia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。