arXiv:2410.21321cs.SIcs.AI2024-10被引 28

结合用户历史与社交上下文,提升低资源印地语暴力内容检测效果

User-Aware Multilingual Abusive Content Detection in Social Media

  • 分模块学习文本与社交上下文特征,融合后用于判断
  • 在SCIDN和MACI数据集上F1得分分别提升4.08%和9.52%
  • 适合关注多语言社交媒体安全的开发者与研究者

尽管已有大量努力遏制社交媒体上的不当内容,但多语言环境使问题更加复杂。低资源语言因数据匮乏而面临更大挑战。本文提出一种针对多种低资源印地语的暴力内容检测新方法。观察发现,一条帖子引发辱骂评论的可能性,以及用户历史与社交上下文特征,对检测具有显著帮助。所提方法首先在两个独立模块中学习社交与文本上下文特征,再融合生成综合表示用于最终预测。在包含150万条(SCIDN)和66.5万条(MACI)多语言评论的数据集上,与经典及前沿方法相比,本方法在SCIDN和MACI数据集上的平均F1得分分别提升4.08%和9.52%。

原文摘要 · Abstract (English)

Despite growing efforts to halt distasteful content on social media, multilingualism has added a new dimension to this problem. The scarcity of resources makes the challenge even greater when it comes to low-resource languages. This work focuses on providing a novel method for abusive content detection in multiple low-resource Indic languages. Our observation indicates that a post's tendency to attract abusive comments, as well as features such as user history and social context, significantly aid in the detection of abusive content. The proposed method first learns social and text context features in two separate modules. The integrated representation from these modules is learned and used for the final prediction. To evaluate the performance of our method against different classical and state-of-the-art methods, we have performed extensive experiments on SCIDN and MACI datasets consisting of 1.5M and 665K multilingual comments, respectively. Our proposed method outperforms state-of-the-art baseline methods with an average increase of 4.08% and 9.52% in F1-scores on SCIDN and MACI datasets, respectively.

暴力内容检测多语言低资源语言社交上下文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。