arXiv:2506.03916cs.CL2025-06EMNLP被引 2

通过合成数据提升仇恨言论检测模型的组合泛化能力

Compositional Generalisation for Explainable Hate Speech Detection

  • 构建合成数据集U-PLEAD,均衡表达在各类语境中出现频率
  • 在8000条人工验证样本上实现优于SOTA的组合泛化性能
  • 适合关注可解释性与鲁棒性仇恨言论检测的研究者

仇恨言论检测对在线内容管理至关重要,但现有模型难以泛化到训练数据之外。这与数据集偏差及使用句子级标签有关,后者未能教会模型仇恨言论的内在结构。本文发现,即使采用更细粒度的跨度级标注(如“艺术家”为攻击目标,“是寄生虫”为去人性化类比),模型仍难以将标签含义与上下文分离,导致训练中未见的表达组合难以识别。为此,我们创建了约36.4万条合成帖子的U-PLEAD数据集,并设计了一个包含约8000条人工验证帖子的新组合泛化基准。结合真实数据与U-PLEAD训练,显著提升了模型在组合泛化上的表现,同时在人工标注的PLEAD数据集上达到当前最优性能。

原文摘要 · Abstract (English)

Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data. This has been linked to dataset biases and the use of sentence-level labels, which fail to teach models the underlying structure of hate speech. In this work, we show that even when models are trained with more fine-grained, span-level annotations (e.g., "artists" is labeled as target and "are parasites" as dehumanising comparison), they struggle to disentangle the meaning of these labels from the surrounding context. As a result, combinations of expressions that deviate from those seen during training remain particularly difficult for models to detect. We investigate whether training on a dataset where expressions occur with equal frequency across all contexts can improve generalisation. To this end, we create U-PLEAD, a dataset of ~364,000 synthetic posts, along with a novel compositional generalisation benchmark of ~8,000 manually validated posts. Training on a combination of U-PLEAD and real data improves compositional generalisation while achieving state-of-the-art performance on the human-sourced PLEAD.

仇恨言论检测组合泛化合成数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。