arXiv:2511.16689cs.CLcs.AI2025-11被引 1

通过概念梯度分析毒性文本中的错误归因,提升检测模型的可解释性。

Concept-Based Interpretability for Toxicity Detection

  • 用概念梯度法衡量语义概念对输出的影响,实现更因果的解释。
  • 发现特定毒词汇过度关联导致误判,通过WCA分数量化这种偏差。
  • 提出无词典增强策略,验证模型是否仍依赖有毒词汇模式。

社交媒体的兴起促进了交流,但也加速了有害内容的传播。尽管文本毒性检测已取得进展,但基于概念的解释研究仍有限。本文利用毒性数据集中多种子类型属性(如脏话、威胁、侮辱、身份攻击、性暗示)作为概念,以识别语言是否具有毒性。然而,概念在目标类别上的不均衡归因常导致分类错误。为此,本文引入基于概念梯度(Concept Gradient, CG)的可解释性方法,通过测量概念变化对模型输出的影响,提供更因果的解释。该方法扩展了传统仅关注输入特征的梯度方法。我们构建了目标词典集(Targeted Lexicon Set),捕捉导致误分类的毒词汇,并计算词-概念对齐(Word-Concept Alignment, WCA)得分,量化这些词汇因过度关联毒概念而引发错误的程度。最后,提出一种无需预设词典的增强策略,生成不包含已有毒词汇的毒性样本,以检验当显式词汇重叠被移除后,模型是否仍存在过度归因现象,从而揭示模型对更广泛毒性语言模式的依赖。

原文摘要 · Abstract (English)

The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic language in textual data, the exploration of concept-based explanations in toxicity detection remains limited. In this study, we leverage various subtype attributes present in toxicity detection datasets, such as obscene, threat, insult, identity attack, and sexual explicit as concepts that serve as strong indicators to identify whether language is toxic. However, disproportionate attribution of concepts towards the target class often results in classification errors. Our work introduces an interpretability technique based on the Concept Gradient (CG) method which provides a more causal interpretation by measuring how changes in concepts directly affect the output of the model. This is an extension of traditional gradient-based methods in machine learning, which often focus solely on input features. We propose the curation of Targeted Lexicon Set, which captures toxic words that contribute to misclassifications in text classification models. To assess the significance of these lexicon sets in misclassification, we compute Word-Concept Alignment (WCA) scores, which quantify the extent to which these words lead to errors due to over-attribution to toxic concepts. Finally, we introduce a lexicon-free augmentation strategy by generating toxic samples that exclude predefined toxic lexicon sets. This approach allows us to examine whether over-attribution persists when explicit lexical overlap is removed, providing insights into the model's attribution on broader toxic language patterns.

可解释性毒性检测概念梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。