用不确定性引导筛选高质量数据,少样本下也能高效识别网络暴力内容。
U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario
- 结合贝叶斯神经网络与主动学习,自动挑选高置信度伪标签用于训练。
- 5样本场景下比基础模型提升14.92%,在不平衡和跨领域数据中表现稳健。
- 兼容多种预训练模型,适合需快速部署的网络内容安全系统。
随着社交媒体普及,用户生成内容激增,其中仇恨、辱骂、攻击性或网络欺凌等行为构成网络暴力,威胁在线生态安全。尽管人工审核仍为主流,但内容量庞大且对审核者心理压力大,亟需自动化检测。现有方法多依赖大规模标注数据,但实际获取成本高、难度大。为此,我们提出面向少样本场景的不确定性引导防火墙U-GIFT,通过自训练提升检测性能。U-GIFT结合主动学习与贝叶斯神经网络(BNNs),基于模型预测的不确定性估计,自动从无标签数据中识别高质量样本,优先选择高置信度伪标签进行训练。大量实验表明,在少样本场景下,U-GIFT显著优于对比基线。在5样本设置下,性能较基础模型提升14.92%。更重要的是,U-GIFT具有良好的用户友好性与可扩展性,适用于多种预训练语言模型(PLMs),在样本不均衡和跨领域场景中表现稳健,展现出强泛化能力。我们认为,U-GIFT为少样本网络暴力内容检测提供了高效方案,有力支持网络内容自动化治理,助力网络安全发展。
原文摘要 · Abstract (English)
With the widespread use of social media, user-generated content has surged on online platforms. When such content includes hateful, abusive, offensive, or cyberbullying behavior, it is classified as toxic speech, posing a significant threat to the online ecosystem's integrity and safety. While manual content moderation is still prevalent, the overwhelming volume of content and the psychological strain on human moderators underscore the need for automated toxic speech detection. Previously proposed detection methods often rely on large annotated datasets; however, acquiring such datasets is both costly and challenging in practice. To address this issue, we propose an uncertainty-guided firewall for toxic speech in few-shot scenarios, U-GIFT, that utilizes self-training to enhance detection performance even when labeled data is limited. Specifically, U-GIFT combines active learning with Bayesian Neural Networks (BNNs) to automatically identify high-quality samples from unlabeled data, prioritizing the selection of pseudo-labels with higher confidence for training based on uncertainty estimates derived from model predictions. Extensive experiments demonstrate that U-GIFT significantly outperforms competitive baselines in few-shot detection scenarios. In the 5-shot setting, it achieves a 14.92\% performance improvement over the basic model. Importantly, U-GIFT is user-friendly and adaptable to various pre-trained language models (PLMs). It also exhibits robust performance in scenarios with sample imbalance and cross-domain settings, while showcasing strong generalization across various language applications. We believe that U-GIFT provides an efficient solution for few-shot toxic speech detection, offering substantial support for automated content moderation in cyberspace, thereby acting as a firewall to promote advancements in cybersecurity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。