arXiv:2508.04913cs.LGcs.CL2025-08中稿 · the Deviant Dynami…被引 6

用120万条数据测试了Transformer模型,发现ELECTRA在仇恨言论检测中表现最佳。

Advancing Hate Speech Detection with Transformers: Insights from the MetaHate

  • 基于MetaHate数据集,对比多种Transformer模型
  • ELECTRA微调后F1得分达0.898,是目前最优结果
  • 揭示讽刺、隐语和标签噪声是主要误判原因

仇恨言论是广泛存在的有害网络言论,包含侮辱性词汇和诽谤内容,对目标个体与群体造成严重的社会、心理甚至身体伤害。随着X(原推特)、脸书、Instagram、Reddit等社交平台持续促进信息传播,仇恨言论也日益成为现实仇恨犯罪的诱因。为此,亟需发展鲁棒的自动化检测方法以应对多样化的社交媒体环境。尽管传统深度学习模型如RNN、LSTM和CNN已取得良好效果,但受限于长程依赖和并行效率问题。本研究基于包含120万条样本的MetaHate数据集——一个由36个数据集组成的元集合,全面评估了BERT、RoBERTa、GPT-2和ELECTRA等主流Transformer模型。结果显示,微调后的ELECTRA表现最佳,F1分数达到0.898。同时,分类错误分析揭示了讽刺、编码语言及标签噪声带来的挑战。

原文摘要 · Abstract (English)

Hate speech is a widespread and harmful form of online discourse, encompassing slurs and defamatory posts that can have serious social, psychological, and sometimes physical impacts on targeted individuals and communities. As social media platforms such as X (formerly Twitter), Facebook, Instagram, Reddit, and others continue to facilitate widespread communication, they also become breeding grounds for hate speech, which has increasingly been linked to real-world hate crimes. Addressing this issue requires the development of robust automated methods to detect hate speech in diverse social media environments. Deep learning approaches, such as vanilla recurrent neural networks (RNNs), long short-term memory (LSTM), and convolutional neural networks (CNNs), have achieved good results, but are often limited by issues such as long-term dependencies and inefficient parallelization. This study represents the comprehensive exploration of transformer-based models for hate speech detection using the MetaHate dataset--a meta-collection of 36 datasets with 1.2 million social media samples. We evaluate multiple state-of-the-art transformer models, including BERT, RoBERTa, GPT-2, and ELECTRA, with fine-tuned ELECTRA achieving the highest performance (F1 score: 0.8980). We also analyze classification errors, revealing challenges with sarcasm, coded language, and label noise.

仇恨言论检测TransformerELECTRAMetaHate

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。