arXiv:2409.16673cs.CLcs.LG2024-09被引 25

SWE2通过融合词与子词信息,提升仇恨言论检测准确率与抗干扰能力。

SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection

  • 结合词级语义与子词知识,增强文本表征能力
  • 无攻击时达0.975准确率与0.953宏F1,优于7个基准模型
  • 在50%消息被恶意篡改下仍保持高鲁棒性,适合实际部署

在线社交网络中的仇恨言论检测近年来成为热门研究方向。由于其广泛传播和快速扩散,仇恨言论加剧偏见并伤害他人,引发产业界与学术界的广泛关注。本文提出一种名为SWE2的新框架,仅依赖消息内容自动识别仇恨言论。该框架同时利用词级语义信息与子词知识,在有无字符级对抗攻击的条件下均表现良好。实验结果表明,无对抗攻击时,模型达到0.975准确率和0.953宏F1,超越7个先进基线模型;在极端对抗攻击(50%消息被篡改)下,仍保持0.967准确率和0.934宏F1,展现出显著鲁棒性。

原文摘要 · Abstract (English)

Hate speech detection on online social networks has become one of the emerging hot topics in recent years. With the broad spread and fast propagation speed across online social networks, hate speech makes significant impacts on society by increasing prejudice and hurting people. Therefore, there are aroused attention and concern from both industry and academia. In this paper, we address the hate speech problem and propose a novel hate speech detection framework called SWE2, which only relies on the content of messages and automatically identifies hate speech. In particular, our framework exploits both word-level semantic information and sub-word knowledge. It is intuitively persuasive and also practically performs well under a situation with/without character-level adversarial attack. Experimental results show that our proposed model achieves 0.975 accuracy and 0.953 macro F1, outperforming 7 state-of-the-art baselines under no adversarial attack. Our model robustly and significantly performed well under extreme adversarial attack (manipulation of 50% messages), achieving 0.967 accuracy and 0.934 macro F1.

仇恨言论文本检测鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。