让多语言仇恨言论检测模型学会人类解释,提升准确与透明度。
Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales

- 训练时对齐模型推理与人工标注理由,增强可解释性。
- 在英文HateXplain和印地语混合语BullySent上提升F分数。
- 能捕捉隐含反穆斯林仇恨的文化特定线索,适合内容审核场景。
在线针对穆斯林群体的仇恨言论常以文化编码、多语言形式出现,传统AI监管系统虽准确但缺乏透明度,易产生偏见、过度删帖或漏判,尤其在脱离社会文化背景时。本文提出一种训练时可解释性框架,使模型推理与人工标注的理由对齐,同时提升分类性能与可解释性。在英文HateXplain和印地语混合语BullySent数据集上进行评估,采用LIME、积分梯度、梯度×输入和注意力机制,分析准确性、解释质量及跨方法一致性。结果表明,基于梯度与注意力的正则化显著提升F分数,增强解释的合理性与忠实度,并有效识别隐含反穆斯林仇恨的文化特征,为多语言、文化敏感的内容审核提供可行路径。
原文摘要 · Abstract (English)
Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。