arXiv:2601.13547cs.CLcs.AI2026-01Conference of the …被引 2

为仇恨言论解释的推理质量设计评估工具,让AI判断更透明可信。

HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations

  • 构建四维指标体系,评估解释中结论清晰度与逻辑一致性。
  • 在六个数据集上验证,发现传统指标无法察觉的解释缺陷。
  • 适合内容审核系统开发者和伦理审查人员使用。

仇恨言论检测是内容审核的关键,但现有评估框架很少考察模型为何判定某文本为仇恨言论。我们提出 extsf{HateXScore},一个包含四个维度的评估套件,用于衡量模型解释的推理质量:(i) 结论明确性,(ii) 引用片段的忠实性与因果依据,(iii) 受保护群体识别(可配置),(iv) 各要素间的逻辑一致性。在六个多样化的仇恨言论数据集上进行评估, extsf{HateXScore} 能揭示标准指标如准确率或 F1 无法捕捉的可解释性缺陷和标注不一致问题。此外,人工评估显示与 extsf{HateXScore} 结果高度一致,验证其作为可信、透明审核工具的实用性。注意:本文包含可能令人不适的敏感内容。

原文摘要 · Abstract (English)

Hateful speech detection is a key component of content moderation, yet current evaluation frameworks rarely assess why a text is deemed hateful. We introduce \textsf{HateXScore}, a four-component metric suite designed to evaluate the reasoning quality of model explanations. It assesses (i) conclusion explicitness, (ii) faithfulness and causal grounding of quoted spans, (iii) protected group identification (policy-configurable), and (iv) logical consistency among these elements. Evaluated on six diverse hate speech datasets, \textsf{HateXScore} is intended as a diagnostic complement to reveal interpretability failures and annotation inconsistencies that are invisible to standard metrics like Accuracy or F1. Moreover, human evaluation shows strong agreement with \textsf{HateXScore}, validating it as a practical tool for trustworthy and transparent moderation. \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}

仇恨言论可解释性评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。