arXiv:2602.13455cs.CLcs.AI2026-02中稿 · IJCAI

用机器学习提升斯瓦希里语隐匿暴力词检测,守护儿童网络安全

Using Machine Learning to Enhance the Detection of Obfuscated Abusive Words in Swahili: A Focus on Child Safety

  • 采用SVM、逻辑回归等模型,结合SMOTE处理数据不平衡问题
  • 在小规模数据集上实现可接受的精准率与召回率,但泛化能力受限
  • 为低资源语言的网络欺凌检测提供方法参考,适合安全与AI交叉研究者

数字技术普及加剧了网络欺凌风险,亟需强化检测与防范措施,尤其针对儿童。本研究聚焦斯瓦希里语中被伪装的暴力语言检测,该语言因资源稀缺面临独特挑战。斯瓦希里语是非洲使用最广泛的语言,母语者超1600万,总使用者逾1亿,遍布东非及中东部分地区。我们采用支持向量机(SVM)、逻辑回归和决策树等机器学习模型,通过参数调优及合成少数类过采样技术(SMOTE)应对数据不平衡。分析显示,尽管模型在高维文本数据中表现良好,但数据集规模小且分布不均限制了结果的可推广性。精确率、召回率与F1分数被深入评估,揭示各模型在识别伪装语言上的差异。研究为构建更安全的儿童网络环境提供支持,呼吁扩充数据集并应用先进机器学习技术。未来工作将关注数据增强、迁移学习及多模态融合,以开发更具文化敏感性的检测机制。

原文摘要 · Abstract (English)

The rise of digital technology has dramatically increased the potential for cyberbullying and online abuse, necessitating enhanced measures for detection and prevention, especially among children. This study focuses on detecting abusive obfuscated language in Swahili, a low-resource language that poses unique challenges due to its limited linguistic resources and technological support. Swahili is chosen due to its popularity and being the most widely spoken language in Africa, with over 16 million native speakers and upwards of 100 million speakers in total, spanning regions in East Africa and some parts of the Middle East. We employed machine learning models including Support Vector Machines (SVM), Logistic Regression, and Decision Trees, optimized through rigorous parameter tuning and techniques like Synthetic Minority Over-sampling Technique (SMOTE) to handle data imbalance. Our analysis revealed that, while these models perform well in high-dimensional textual data, our dataset's small size and imbalance limit our findings' generalizability. Precision, recall, and F1 scores were thoroughly analyzed, highlighting the nuanced performance of each model in detecting obfuscated language. This research contributes to the broader discourse on ensuring safer online environments for children, advocating for expanded datasets and advanced machine-learning techniques to improve the effectiveness of cyberbullying detection systems. Future work will focus on enhancing data robustness, exploring transfer learning, and integrating multimodal data to create more comprehensive and culturally sensitive detection mechanisms.

网络安全低资源语言机器学习儿童保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。