系统梳理英文文本仇恨言论多标签分类的模型与数据集,揭示研究现状与挑战。
A Survey of Machine Learning Models and Datasets for the Multi-label Classification of Textual Hate Speech in English
- 首次系统综述46篇相关文献,归纳28个可用数据集
- 发现多数模型依赖BERT和RNN,评估标准不统一
- 指出数据不平衡、标注质量差等关键问题,适合研究者参考
网络仇恨言论的传播对个人、在线社区及社会具有严重负面影响。海量仇恨内容促使内容审核与执法机构、研究人员关注机器学习模型在自动识别仇恨言论中的应用。尽管多数研究将仇恨言论分类视为二分类任务,但实际应用常需区分目标、严重程度或合法性等子类型,且这些类别可能重叠。为此,研究者构建了面向多标签分类的文本仇恨言论数据集与模型。本文首次系统性综述了该领域46篇英文文献,总结28个适用于训练多标签分类模型的数据集,揭示其在标签集合、规模、元概念、标注流程及标注者一致性方面存在显著异质性。对24项提出适用模型的研究分析显示,评估方法不一致,且偏好使用基于BERT和循环神经网络(RNN)的架构。我们识别出训练数据不平衡、过度依赖众包平台、数据集小而稀疏、方法论缺乏统一等关键开放问题,并提出十项研究建议。
原文摘要 · Abstract (English)
The dissemination of online hate speech can have serious negative consequences for individuals, online communities, and entire societies. This and the large volume of hateful online content prompted both practitioners', i.e., in content moderation or law enforcement, and researchers' interest in machine learning models to automatically classify instances of hate speech. Whereas most scientific works address hate speech classification as a binary task, practice often requires a differentiation into sub-types, e.g., according to target, severity, or legality, which may overlap for individual content. Hence, researchers created datasets and machine learning models that approach hate speech classification in textual data as a multi-label problem. This work presents the first systematic and comprehensive survey of scientific literature on this emerging research landscape in English (N=46). We contribute with a concise overview of 28 datasets suited for training multi-label classification models that reveals significant heterogeneity regarding label-set, size, meta-concept, annotation process, and inter-annotator agreement. Our analysis of 24 publications proposing suitable classification models further establishes inconsistency in evaluation and a preference for architectures based on Bidirectional Encoder Representation from Transformers (BERT) and Recurrent Neural Networks (RNNs). We identify imbalanced training data, reliance on crowdsourcing platforms, small and sparse datasets, and missing methodological alignment as critical open issues and formulate ten recommendations for research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。