用大模型+流式学习实时检测网络欺凌,效果接近90%且可解释。
Promoting Security and Trust on Social Networks: Explainable Cyberbullying Detection Using Large Language Models in a Stream-Based Machine Learning Framework
- 结合大语言模型做特征工程,流式学习逐条处理新数据。
- 实验中各项指标接近90%,优于已有方法。
- 提供可解释仪表盘,适合平台安全团队使用。
社交媒体平台虽促进即时沟通,但也催生网络欺凌等负面行为。本文提出一种基于流式机器学习与大语言模型的实时检测方案,利用LLM进行特征工程,支持增量处理新样本,应对网络暴力内容的动态演化。系统配备可解释性仪表盘,提升可信度与问责性。实验结果显示,在所有评估指标上表现接近90%,超越现有文献中的对比方法。该方案有助于及时发现恶意行为,防止长期骚扰,降低社会负面影响。
原文摘要 · Abstract (English)
Social media platforms enable instant and ubiquitous connectivity and are essential to social interaction and communication in our technological society. Apart from its advantages, these platforms have given rise to negative behaviors in the online community, the so-called cyberbullying. Despite the many works involving generative Artificial Intelligence (AI) in the literature lately, there remain opportunities to study its performance apart from zero/few-shot learning strategies. Accordingly, we propose an innovative and real-time solution for cyberbullying detection that leverages stream-based Machine Learning (ML) models able to process the incoming samples incrementally and Large Language Models (LLMS) for feature engineering to address the evolving nature of abusive and hate speech online. An explainability dashboard is provided to promote the system's trustworthiness, reliability, and accountability. Results on experimental data report promising performance close to 90 % in all evaluation metrics and surpassing those obtained by competing works in the literature. Ultimately, our proposal contributes to the safety of online communities by timely detecting abusive behavior to prevent long-lasting harassment and reduce the negative consequences in society.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。