探索仇恨言论检测中人类解释的差异性,重构评估框架。
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection

- 统一多种模型与评估指标,系统测试不同标签和解释表示方式
- 软标签和软解释更优,能更好捕捉人类判断的多样性
- 为主观任务提供更合理的评价思路,适合研究可解释AI的学者
人类在标注过程中存在普遍分歧,但基于词粒度的人类解释差异仍鲜被研究。现有评估方法难以应对这种分歧,尤其在如何聚合解释(超越多数投票)方面尚不明确。在仇恨言论检测这类主观性任务中,解释可能反映不同推理风格、价值观与理解方式。本文通过统一协议,系统重实现多种模型、训练策略、损失函数及评估指标,涵盖不同的标签与解释表示空间。分类评估聚焦预测性和分布性,可解释性评估从合理性、忠实度与复杂度三个维度展开。结果表明,软标签与软解释表现更优,凸显其对人类判断多样性的捕捉能力,提示需重新思考主观自然语言处理任务的评估范式。
原文摘要 · Abstract (English)
Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains far less explored. At the same time, it is unclear how to best evaluate human labels and rationales -- or even how to best aggregate rationales beyond majority vote -- in light of this variation. Yet, rationales may provide additional insights into the richness of human reasoning, that may differ in style, values and interpretations -- especially in subjective NLP tasks like hate speech detection. In this work, we unify diverse models, training strategies, loss functions, and existing evaluation metrics under a single protocol by systematically re-implementing them across different label and rationale representation spaces. Classification metrics are organized around two key properties -- predictive and distributional -- while explainability metrics through three complementary dimensions: plausibility, faithfulness, and complexity. In this unified supervision framework, we evaluate model behavior across classification and explainability metrics, as well as metric sensitivity to the choice of label (hard and soft) and rationale representation space (hard, intermediate and soft). Results show that both hard and soft metrics favor softer representations, highlighting their effectiveness in capturing variation and the need to rethink evaluation in subjective NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。