arXiv:2601.13018cs.CLcs.AI2026-01中稿 · "EAI AFRICOMM 2025…被引 2

提出双向注意力模型,提升仇恨言论检测的解释性与稳定性。

Bi-Attention HateXplain : Taking into account the sequential aspect of data during explainability in a multi-task context

  • 引入双向RNN捕捉文本序列特性,增强解释一致性。
  • 在HateXplain数据集上提升检测准确率并减少偏差。
  • 适合关注可解释性和公平性的自然语言处理研究者。

互联网和在线社交网络的发展带来了诸多益处,但也加剧了仇恨言论的泛滥,成为全球主要威胁。为提高黑箱模型在仇恨言论检测中的可靠性,后处理解释方法如LIME、SHAP和LRP在模型训练后提供解释。相比之下,基于HateXplain基准的多任务方法同时学习解释与分类。然而,现有方法在预测注意力时波动显著,导致解释不一致、预测不稳定及学习困难。为此,本文提出BiAtt-BiRNN-HateXplain模型,通过双向RNN层在解释过程中考虑输入数据的序列特性,提升可解释性,并借助多任务学习(解释与分类)实现更优分类性能,减少对特定群体的无意识偏差。实验结果表明,该模型在HateXplain数据集上显著提升了检测性能、解释质量并降低了偏差。

原文摘要 · Abstract (English)

Technological advances in the Internet and online social networks have brought many benefits to humanity. At the same time, this growth has led to an increase in hate speech, the main global threat. To improve the reliability of black-box models used for hate speech detection, post-hoc approaches such as LIME, SHAP, and LRP provide the explanation after training the classification model. In contrast, multi-task approaches based on the HateXplain benchmark learn to explain and classify simultaneously. However, results from HateXplain-based algorithms show that predicted attention varies considerably when it should be constant. This attention variability can lead to inconsistent interpretations, instability of predictions, and learning difficulties. To solve this problem, we propose the BiAtt-BiRNN-HateXplain (Bidirectional Attention BiRNN HateXplain) model which is easier to explain compared to LLMs which are more complex in view of the need for transparency, and will take into account the sequential aspect of the input data during explainability thanks to a BiRNN layer. Thus, if the explanation is correctly estimated, thanks to multi-task learning (explainability and classification task), the model could classify better and commit fewer unintentional bias errors related to communities. The experimental results on HateXplain data show a clear improvement in detection performance, explainability and a reduction in unintentional bias.

仇恨言论可解释性多任务学习序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。