通过上下文限定的难例挖掘,提升隐性仇恨言论跨数据集检测能力。
Aligning Implied Statements for Implicit Hate Speech Generalizability with Context-Bounded Semi-hard Negative Mining

- 用三元组对齐含蓄表述,聚焦近似混淆样本学习
- 在IHC/SBIC/DynaHate上跨域准确率提升1.5%-3.2%
- 适合需鲁棒泛化性的仇恨言论检测研究者
隐性仇恨言论分类仍具挑战,因意图常以暗示和语境隐藏而非明确辱骂。现有监督对比方法虽在本域表现良好,但易过拟合表面线索,跨数据集迁移能力差。本文提出ImpSH框架,基于三元组对齐有隐含表述的文本,并采用上下文限定的半难负样本挖掘,使学习聚焦于近似混淆样本。同时考察通过数据增强生成正样本的AugSH方法。在IHC、SBIC与DynaHate数据集上,使用BERT与HateBERT进行受控评估,ImpSH是标准监督对比基线的可行替代方案,且在匹配预处理与调参预算下常显著提升跨域性能。表示分析显示正样本对更紧密,全局分布更均衡;定性最近邻案例揭示领域迁移下的典型误检。结果表明,通过上下文限定挖掘对齐隐含表述,可建立更稳定、类双射的映射关系,克服传统聚类式表征学习的不稳定性。
原文摘要 · Abstract (English)
Classifying implicit hate speech remains a challenge, as intent is often masked through insinuation and context rather than explicit slurs. Prior supervised contrastive approaches improve in-domain detection but can overfit surface cues and struggle to transfer across datasets. We propose ImpSH, a triplet-based framework that aligns posts with implied statements when available and uses context-bounded semi-hard negatives to focus learning on near confusions. We also examine AugSH, which forms positives via data augmentation. In controlled evaluations on IHC, SBIC, and DynaHate with BERT and HateBERT, ImpSH is a viable alternative to standard supervised contrastive baselines and often improves cross-domain performance under matched preprocessing and tuning budgets. Representation analysis using alignment and uniformity indicates tighter positive pairs with balanced global spread, and qualitative nearest-neighbor case studies illustrate typical false negatives under domain shift. These results demonstrate that aligning posts with their implied statements via context-bounded mining provides a more stable, bijective-like mapping to related insinuations, overcoming the volatility inherent in traditional clustering-based representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。