arXiv:2506.02753cs.CL2025-06被引 3

用多任务+主动学习提升阿拉伯语辱骂内容检测精度

Multi-task Learning with Active Learning for Arabic Offensive Speech Detection

  • 联合暴力、低俗任务共享表征,动态调整任务权重
  • 仅用少量标注样本即达85.42%宏F1,优于现有方法
  • 适合数据稀缺的阿拉伯语内容安全场景

社交媒体的快速发展加剧了辱骂、暴力和粗俗言论的传播,对社会与网络安全构成严重威胁。阿拉伯语文本的检测尤为复杂,受限于标注数据少、方言差异大及语言本身复杂性。本文提出一种融合多任务学习(MTL)与主动学习的新框架,通过联合训练暴力和低俗两个辅助任务,利用共享表征提升辱骂内容检测准确率。训练中动态调整任务权重以平衡各任务贡献。针对标注数据稀缺问题,采用多种不确定性采样策略主动选择最具信息量的样本进行迭代训练,并引入加权表情符号处理以更好捕捉语义线索。在OSACT2022数据集上的实验表明,该框架达到85.42%的宏F1分数,显著优于现有方法,且所需微调样本更少。研究验证了MTL与主动学习结合在资源受限环境下高效精准检测攻击性语言的潜力。

原文摘要 · Abstract (English)

The rapid growth of social media has amplified the spread of offensive, violent, and vulgar speech, which poses serious societal and cybersecurity concerns. Detecting such content in Arabic text is particularly complex due to limited labeled data, dialectal variations, and the language's inherent complexity. This paper proposes a novel framework that integrates multi-task learning (MTL) with active learning to enhance offensive speech detection in Arabic social media text. By jointly training on two auxiliary tasks, violent and vulgar speech, the model leverages shared representations to improve the detection accuracy of the offensive speech. Our approach dynamically adjusts task weights during training to balance the contribution of each task and optimize performance. To address the scarcity of labeled data, we employ an active learning strategy through several uncertainty sampling techniques to iteratively select the most informative samples for model training. We also introduce weighted emoji handling to better capture semantic cues. Experimental results on the OSACT2022 dataset show that the proposed framework achieves a state-of-the-art macro F1-score of 85.42%, outperforming existing methods while using significantly fewer fine-tuning samples. The findings of this study highlight the potential of integrating MTL with active learning for efficient and accurate offensive language detection in resource-constrained settings.

多任务学习主动学习阿拉伯语内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。