用可解释性分析模型为何误判网络有害内容,揭示隐藏缺陷。
Beyond Accuracy: An Explainability-Driven Analysis of Harmful Content Detection
- 用Shapley和Integrated Gradients分析RoBERTa模型决策依据。
- 准确率94%,AUC达0.93,但仍有误判与漏判。
- 适合关注AI透明度与内容审核可信度的研究者。
尽管自动化有害内容检测系统常用于监控在线平台,但审核员和用户往往无法理解其预测逻辑。现有研究多聚焦提升分类准确率,却忽视了神经模型在边缘、语境及政治敏感情境下判断有害内容的原因。本文基于Civil Comments数据集训练的RoBERTa分类器,采用Shapley Additive Explanations(SHAP)与Integrated Gradients两种后处理解释方法,分析模型在正确预测与系统性失败案例中的行为。尽管整体表现优异,AUC为0.93,准确率为0.94,但可解释性分析揭示了聚合指标无法反映的局限:Integrated Gradients产生更分散的上下文归因,而SHAP则更聚焦于显式词汇线索。二者输出差异导致假阳性和假阴性。定性案例显示常见失败模式包括间接攻击、词汇过度归因与政治话语误判。结果表明,可解释AI能通过暴露模型不确定性,增强人工介入审核的可信度,更重要的是,它应作为透明度与诊断工具,而非性能提升手段。
原文摘要 · Abstract (English)
Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on increasing classification accuracy, little focus has been placed on comprehending why neural models identify content as harmful, especially when it comes to borderline, contextual, and politically sensitive situations. In this work, a neural harmful content detection model trained on the Civil Comments dataset is analyzed explainability-drivenly. Two popular post-hoc explanation methods, Shapley Additive Explanations and Integrated Gradients, are used to analyze the behavior of a RoBERTa-based classifier in both correct predictions and systematic failure cases. Despite strong overall performance, with an area under the curve of 0.93 and an accuracy of 0.94, the analysis reveals limitations that are not observable from aggregate evaluation metrics alone. Integrated Gradients appear to extract more diffuse contextual attributions while Shapley Additive Explanations extract more focused attributions on explicit lexical cues. The consequent divergence in their outputs manifests in both false negatives and false positives. Qualitative case studies reveal recurring failure modes such as indirect toxicity, lexical over-attribution, or political discourse. The results suggest that explainable AI can foster human-in-the-loop moderation by exposing model uncertainty and increasing the interpretable rationale behind automated decisions. Most importantly, this work highlights the role of explainability as a transparency and diagnostic resource for online harmful content detection systems rather than as a performance-enhancing lever.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。