arXiv:2504.00045cs.CLcs.CY2025-04被引 7

用AI模型检测4chan上仇恨言论,发现11.2%内容含仇恨信息

Measuring Online Hate on 4chan using Pre-trained Deep Learning Models

  • 采用RoBERTa和Detoxify等预训练模型分析匿名论坛内容
  • 11.20%的帖子被识别为包含种族、性别等各类仇恨言论
  • 同时识别毒害性内容并进行主题分析,适合研究网络暴力者

在线仇恨言论会对个人和群体造成伤害,尤其在4chan这类无审核平台,用户可匿名发帖。本研究利用先进的自然语言处理模型(如RoBERTa和Detoxify),对4chan的/pol/版进行深度分析,量化仇恨言论的泛滥程度。通过多类别仇恨分类(如种族主义、性别歧视、宗教仇恨等),以及毒害内容(如身份攻击、威胁)识别和主题建模,揭示了仇恨言论的多样表现形式。结果显示,该数据集中11.20%的内容被识别为含有不同类别的仇恨言论。研究证实了现实网络环境中仇恨言论的复杂性和易变性,为理解非监管平台上的内容治理提供了重要依据。

原文摘要 · Abstract (English)

Online hate speech can harmfully impact individuals and groups, specifically on non-moderated platforms such as 4chan where users can post anonymous content. This work focuses on analysing and measuring the prevalence of online hate on 4chan's politically incorrect board (/pol/) using state-of-the-art Natural Language Processing (NLP) models, specifically transformer-based models such as RoBERTa and Detoxify. By leveraging these advanced models, we provide an in-depth analysis of hate speech dynamics and quantify the extent of online hate non-moderated platforms. The study advances understanding through multi-class classification of hate speech (racism, sexism, religion, etc.), while also incorporating the classification of toxic content (e.g., identity attacks and threats) and a further topic modelling analysis. The results show that 11.20% of this dataset is identified as containing hate in different categories. These evaluations show that online hate is manifested in various forms, confirming the complicated and volatile nature of detection in the wild.

仇恨言论NLP4chanAI检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。