用BERT和FastText提升阿拉伯语网络欺凌检测准确率至98%
Enhanced Arabic-language cyberbullying detection: deep embedding and transformer (BERT) approaches
- 结合BERT与LSTM/Bi-LSTM模型,利用深度嵌入提取语言特征
- 在10,662条推文上测试,最优模型达98%准确率
- 为阿拉伯语网络欺凌检测提供可推广的高精度方案
智能手机与社交媒体平台(如X)的普及使青少年面临网络欺凌威胁。现有自动检测方法多针对英语,阿拉伯语相关研究稀缺。本文构建包含10,662条X平台推文的标注数据集,经Kappa工具验证质量。实验对比多种深度学习模型:先测试LSTM与Bi-LSTM搭配不同词嵌入,再引入预训练双向编码器BERT,最终发现LSTM-BERT与Bi-LSTM-BERT模型达到97%准确率;而使用FastText嵌入的Bi-LSTM模型表现更优,达98%准确率。结果具有泛化性,为阿拉伯语网络欺凌检测提供了有效技术路径。
原文摘要 · Abstract (English)
Recent technological advances in smartphones and communications, including the growth of such online platforms as massive social media networks such as X (formerly known as Twitter) endangers young people and their emotional well-being by exposing them to cyberbullying, taunting, and bullying content. Most proposed approaches for automatically detecting cyberbullying have been developed around the English language, and methods for detecting Arabic-language cyberbullying are scarce. Methods for detecting Arabic-language cyberbullying are especially scarce. This paper aims to enhance the effectiveness of methods for detecting cyberbullying in Arabic-language content. We assembled a dataset of 10,662 X posts, pre-processed the data, and used the kappa tool to verify and enhance the quality of our annotations. We conducted four experiments to test numerous deep learning models for automatically detecting Arabic-language cyberbullying. We first tested a long short-term memory (LSTM) model and a bidirectional long short-term memory (Bi-LSTM) model with several experimental word embeddings. We also tested the LSTM and Bi-LSTM models with a novel pre-trained bidirectional encoder from representations (BERT) and then tested them on a different experimental models BERT again. LSTM-BERT and Bi-LSTM-BERT demonstrated a 97% accuracy. Bi-LSTM with FastText embedding word performed even better, achieving 98% accuracy. As a result, the outcomes are generalize
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。