用道德理由指导注意力,让仇恨言论检测模型自己解释判断依据。
Self-Explaining Hate Speech Detection with Moral Rationales
- 基于道德基础理论,用专家标注的道德理由直接监督模型注意力。
- 在巴西葡萄牙语数据集上,准确率提升且解释更可信,但简洁性下降。
- 适合需要可解释、文化敏感的仇恨言论检测场景。
现有仇恨言论检测模型往往不透明,依赖表面词汇线索,易受虚假相关影响,限制了鲁棒性、可解释性和文化适应性。我们提出监督道德理由注意力(SMRA),首个将道德理由作为直接监督信号以对齐注意力的自解释仇恨言论检测框架。基于道德基础理论,SMRA 将词级注意力与专家标注的道德理由对齐,引导模型关注具有道德意义的文本片段。不同于以往的论据监督或事后解释方法,SMRA 将道德理由监督直接融入训练目标,生成内在可解释且上下文相关的解释。为支持该框架,我们还引入 HateBRMoralXplain,一个标注了仇恨标签、道德类别、词级道德理由及社会政治元数据的巴西葡萄牙语基准数据集。在二分类仇恨言论检测和多标签道德情感分类任务中,SMRA 均持续提升性能,并增强解释的忠实性与合理性。尽管解释变得更简洁,充分性有所下降,表明理由更紧凑且信息量更高。公平性保持稳定,说明解释质量提升未带来显著偏差权衡。
原文摘要 · Abstract (English)
Existing hate speech detection models are often opaque and rely on surface-level lexical cues, which makes them vulnerable to spurious correlations and limits robustness, interpretability and cultural contextualization. We propose Supervised Moral Rationale Attention (SMRA), the first self-explaining hate speech detection framework to incorporate moral rationales as direct supervision for attention alignment. Based on Moral Foundations Theory, SMRA aligns token-level attention with expert-annotated moral rationales, guiding models to attend to morally salient spans. Unlike prior rationale-supervised or post-hoc approaches, SMRA integrates moral rationale supervision directly into the training objective, producing inherently interpretable and contextualized explanations. To support our framework, we also introduce HateBRMoralXplain, a Brazilian Portuguese benchmark dataset annotated with hate labels, moral categories, token-level moral rationales, and socio-political metadata. Across binary hate speech detection and multi-label moral sentiment classification, SMRA consistently improves performance while enhancing both faithful and plausible explanations. Although explanations become more concise, sufficiency decreases, indicating more compact and informative rationales. Fairness remains stable, suggesting that improvements in explanation quality do not introduce significant bias trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。