arXiv:2409.05134cs.CL2024-09被引 1

调整文本预处理顺序并结合集成方法,显著提升社交媒体仇恨内容识别准确率。

Hate Content Detection via Novel Pre-Processing Sequencing and Ensemble Methods

  • 通过实验对比不同预处理步骤的先后顺序,找出最优流程。
  • 在三个公开数据集上达到最高95.14%的分类准确率。
  • 适合关注文本分类优化与网络内容安全的研究者和工程师。

社交媒体,尤其是推特(Twitter),近年来出现了大量如网络欺凌和仇恨言论等事件。因此,识别仇恨言论已成为当务之急。本文提出一种计算框架,用于遏制网络上的仇恨内容。具体而言,该研究系统地探讨了文本预处理操作顺序对仇恨内容识别的影响。最佳预处理序列与支持向量机(SVM)、随机森林(Random Forest)、决策树(Decision Tree)、逻辑回归(Logistic Regression)和K近邻(K-Neighbor)等主流分类器结合后,显著提升了性能。此外,该最优预处理序列还与袋装法(bagging)、提升法(boosting)和堆叠法(stacking)等集成学习方法联合使用,进一步优化结果。采用三个公开基准数据集(WZ-LS、DT、FOUNTA)进行评估,所提方法在测试中达到最高95.14%的准确率,验证了独特预处理流程与集成分类器结合的有效性。

原文摘要 · Abstract (English)

Social media, particularly Twitter, has seen a significant increase in incidents like trolling and hate speech. Thus, identifying hate speech is the need of the hour. This paper introduces a computational framework to curb the hate content on the web. Specifically, this study presents an exhaustive study of pre-processing approaches by studying the impact of changing the sequence of text pre-processing operations for the identification of hate content. The best-performing pre-processing sequence, when implemented with popular classification approaches like Support Vector Machine, Random Forest, Decision Tree, Logistic Regression and K-Neighbor provides a considerable boost in performance. Additionally, the best pre-processing sequence is used in conjunction with different ensemble methods, such as bagging, boosting and stacking to improve the performance further. Three publicly available benchmark datasets (WZ-LS, DT, and FOUNTA), were used to evaluate the proposed approach for hate speech identification. The proposed approach achieves a maximum accuracy of 95.14% highlighting the effectiveness of the unique pre-processing approach along with an ensemble classifier.

仇恨言论识别文本预处理集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。