多语言多标签网络欺凌检测新框架,提升低资源语言识别效果
HMS-BERT: Hybrid Multi-Task Self-Training for Multilingual and Multi-Label Cyberbullying Detection
- 融合上下文与手工特征,双任务联合优化分类
- 自训练机制在低资源语言上实现98.47%多标签F1分数
- 适合跨语言、多类别网络暴力识别场景
社交媒体上的网络欺凌具有多语言性和多维度特征,现有方法常受限于单语假设或单任务设计,在真实多语言多标签场景中表现不佳。本文提出HMS-BERT,一种基于预训练多语言BERT的混合多任务自训练框架,结合上下文表示与手工语言特征,联合优化细粒度多标签辱骂分类和三分类主任务。为缓解低资源语言标注数据稀缺问题,引入基于置信度的迭代自训练策略,促进跨语言知识迁移。在四个公开数据集上的实验表明,HMS-BERT在多标签任务上达到最高0.9847的宏F1得分,在主分类任务上准确率达0.6775。消融实验证实了各组件的有效性。
原文摘要 · Abstract (English)
Cyberbullying on social media is inherently multilingual and multi-faceted, where abusive behaviors often overlap across multiple categories. Existing methods are commonly limited by monolingual assumptions or single-task formulations, which restrict their effectiveness in realistic multilingual and multi-label scenarios. In this paper, we propose HMS-BERT, a hybrid multi-task self-training framework for multilingual and multi-label cyberbullying detection. Built upon a pretrained multilingual BERT backbone, HMS-BERT integrates contextual representations with handcrafted linguistic features and jointly optimizes a fine-grained multi-label abuse classification task and a three-class main classification task. To address labeled data scarcity in low-resource languages, an iterative self-training strategy with confidence-based pseudo-labeling is introduced to facilitate cross-lingual knowledge transfer. Experiments on four public datasets demonstrate that HMS-BERT achieves strong performance, attaining a macro F1-score of up to 0.9847 on the multi-label task and an accuracy of 0.6775 on the main classification task. Ablation studies further verify the effectiveness of the proposed components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。