用因果模型分离仇恨言论的深层意图与表面风格,提升跨风格检测能力
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
- 构建因果图分解仇恨言论的语境、动机、目标和风格等关键因素
- 在多个数据集上超越基线模型,跨平台检测准确率显著提升
- 支持反事实推理,适合需要可解释性与鲁棒性的内容安全场景
网络仇恨言论的泛滥严重威胁网络和谐。显性仇恨言论可通过明显辱骂词识别,而隐性仇恨言论常通过讽刺、反语、刻板印象或隐晦语言表达,难以察觉。现有检测模型主要依赖表层语言特征,难以在不同风格间泛化。此外,不同平台上的仇恨言论针对不同群体且风格各异,可能引入虚假相关性,进一步挑战检测方法。基于此,我们假设仇恨言论生成可建模为包含语境、创作者动机、目标和风格的因果图。受此启发,提出CADET框架,将仇恨言论解耦为可解释的潜在因子,并控制混杂变量,从而分离真实仇恨意图与表面语言特征。此外,通过潜空间干预风格,实现反事实推理,使模型能稳健识别各种形式的仇恨言论。全面实验表明CADET性能优越,验证了因果先验在提升可泛化仇恨言论检测中的潜力。
原文摘要 · Abstract (English)
The proliferation of online hate speech poses a significant threat to the harmony of the web. While explicit hate is easily recognized through overt slurs, implicit hate speech is often conveyed through sarcasm, irony, stereotypes, or coded language -- making it harder to detect. Existing hate speech detection models, which predominantly rely on surface-level linguistic cues, fail to generalize effectively across diverse stylistic variations. Moreover, hate speech spread on different platforms often targets distinct groups and adopts unique styles, potentially inducing spurious correlations between them and labels, further challenging current detection approaches. Motivated by these observations, we hypothesize that the generation of hate speech can be modeled as a causal graph involving key factors: contextual environment, creator motivation, target, and style. Guided by this graph, we propose CADET, a causal representation learning framework that disentangles hate speech into interpretable latent factors and then controls confounders, thereby isolating genuine hate intent from superficial linguistic cues. Furthermore, CADET allows counterfactual reasoning by intervening on style within the latent space, naturally guiding the model to robustly identify hate speech in varying forms. CADET demonstrates superior performance in comprehensive experiments, highlighting the potential of causal priors in advancing generalizable hate speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。