arXiv:2511.07065cs.CLcs.LG2025-11AAAI被引 5

让模型注意力对齐人类解释,提升仇恨言论检测的可解释性与公平性

Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection

  • 在Transformer中引入监督注意力机制,使模型关注点匹配人类标注的理由
  • 在英文和葡萄牙文数据集上,解释能力提升2.4倍,且更贴近人类判断
  • 既增强可解释性,又保持高公平性,适合需透明决策的伦理场景

深度学习模型的不透明性给仇恨言论检测系统的伦理部署带来挑战。为此,我们提出监督理性注意力(SRA)框架,通过显式对齐模型注意力与人类理由,提升分类任务的可解释性与公平性。SRA将监督注意力机制融入基于Transformer的分类器,优化联合目标函数,包含标准分类损失与最小化注意力权重和人工标注理由差异的对齐损失。我们在英语(HateXplain)和葡萄牙语(HateBRXplain)的仇恨言论基准数据集上进行了评估,这些数据集带有理由标注。实验表明,SRA在解释能力上比现有基线高出2.4倍,生成的词级别解释更忠实、更符合人类认知。在公平性方面,SRA在所有指标上表现良好,在识别针对身份群体的有毒内容上达到第二佳表现,其他指标也保持相当水平。结果表明,将人类理由融入注意力机制可在不损害公平性的前提下,显著提升可解释性与可信度。

原文摘要 · Abstract (English)

The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a framework that explicitly aligns model attention with human rationales, improving both interpretability and fairness in hate speech classification. SRA integrates a supervised attention mechanism into transformer-based classifiers, optimizing a joint objective that combines standard classification loss with an alignment loss term that minimizes the discrepancy between attention weights and human-annotated rationales. We evaluated SRA on hate speech benchmarks in English (HateXplain) and Portuguese (HateBRXplain) with rationale annotations. Empirically, SRA achieves 2.4x better explainability compared to current baselines, and produces token-level explanations that are more faithful and human-aligned. In terms of fairness, SRA achieves competitive fairness across all measures, with second-best performance in detecting toxic posts targeting identity groups, while maintaining comparable results on other metrics. These findings demonstrate that incorporating human rationales into attention mechanisms can enhance interpretability and faithfulness without compromising fairness.

可解释性仇恨言论注意力机制公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。