arXiv:2603.09984cs.CL2026-03

融合BERT/CNN/LSTM的模型高效识别网络暴力语言,准确率达99%。

An Efficient Hybrid Deep Learning Approach for Detecting Online Abusive Language

  • 用BERT+CNN+LSTM融合架构捕捉语义、上下文和序列特征
  • 在77,620条恶意文本上实现99%的精准检测性能
  • 适合需要高精度识别网络欺凌内容的平台与监管场景

数字时代使社交媒体和在线论坛得以普及,近45%全球人口可自由表达。然而也加剧了网络骚扰、霸凌及仇恨言论等有害行为,在社交网络、即时通讯和游戏社区中普遍存在。研究显示65%家长注意到敌意在线行为,三分之一青少年在移动游戏中遭遇霸凌。大量滥用内容每日生成并传播,不仅在公开网页,也在暗网论坛中存在。恶意用户常使用特定词汇或编码短语规避检测。为此,我们提出一种融合BERT、CNN与LSTM架构的混合深度学习模型,结合ReLU激活函数,用于检测包括YouTube评论、论坛讨论和暗网帖子在内的多平台滥用语言。该模型在包含77,620条恶意文本与272,214条正常文本(比例1:3.5)的多样化且不平衡数据集上表现优异,各项评估指标(精确率、召回率、准确率、F1分数、AUC)均接近99%。该方法能有效捕捉文本的语义、上下文与序列模式,即使在真实世界中高度偏斜的数据集下仍具鲁棒性。

原文摘要 · Abstract (English)

The digital age has expanded social media and online forums, allowing free expression for nearly 45% of the global population. Yet, it has also fueled online harassment, bullying, and harmful behaviors like hate speech and toxic comments across social networks, messaging apps, and gaming communities. Studies show 65% of parents notice hostile online behavior, and one-third of adolescents in mobile games experience bullying. A substantial volume of abusive content is generated and shared daily, not only on the surface web but also within dark web forums. Creators of abusive comments often employ specific words or coded phrases to evade detection and conceal their intentions. To address these challenges, we propose a hybrid deep learning model that integrates BERT, CNN, and LSTM architectures with a ReLU activation function to detect abusive language across multiple online platforms, including YouTube comments, online forum discussions, and dark web posts. The model demonstrates strong performance on a diverse and imbalanced dataset containing 77,620 abusive and 272,214 non-abusive text samples (ratio 1:3.5), achieving approximately 99% across evaluation metrics such as Precision, Recall, Accuracy, F1-score, and AUC. This approach effectively captures semantic, contextual, and sequential patterns in text, enabling robust detection of abusive content even in highly skewed datasets, as encountered in real-world scenarios.

网络暴力深度学习文本检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。