arXiv:2512.23835cs.CLcs.AI2025-12

对比分析两种新闻偏见检测模型的决策机制,发现其误报原因不同。

Explaining News Bias Detection: A Comparative SHAP Analysis of Transformer Model Decision Mechanisms

  • 用SHAP分析模型对词语的权重分配方式
  • 域适应模型误报率比普通模型低63%
  • 误报主要源于语境模糊而非明确偏见信号

自动化新闻偏见检测广泛用于支持新闻分析与媒体问责,但对其决策过程和失败原因了解甚少。本文对比分析两种基于Transformer的偏见检测模型:在BABE数据集上微调的偏见检测器和在相同数据集上微调的领域自适应预训练RoBERTa模型,采用基于SHAP的解释方法。通过分析正确与错误预测中的词级归因,揭示不同模型如何操作语言偏见信号。结果显示,尽管两模型均关注类似评价性语言类别,但在信号整合方式上存在显著差异。偏见检测器对假阳性赋予更强内部证据,导致对中立新闻内容系统性误标;而域适应模型的归因模式更贴近预测结果,误报率降低63%。进一步表明,模型错误源于不同语言机制,假阳性由话语层面模糊性驱动,而非显式偏见线索。研究强调了可解释性评估在偏见检测系统中的重要性,表明架构与训练策略显著影响模型可靠性与新闻场景下的适用性。

原文摘要 · Abstract (English)

Automated bias detection in news text is heavily used to support journalistic analysis and media accountability, yet little is known about how bias detection models arrive at their decisions or why they fail. In this work, we present a comparative interpretability study of two transformer-based bias detection models: a bias detector fine-tuned on the BABE dataset and a domain-adapted pre-trained RoBERTa model fine-tuned on the BABE dataset, using SHAP-based explanations. We analyze word-level attributions across correct and incorrect predictions to characterize how different model architectures operationalize linguistic bias. Our results show that although both models attend to similar categories of evaluative language, they differ substantially in how these signals are integrated into predictions. The bias detector model assigns stronger internal evidence to false positives than to true positives, indicating a misalignment between attribution strength and prediction correctness and contributing to systematic over-flagging of neutral journalistic content. In contrast, the domain-adaptive model exhibits attribution patterns that better align with prediction outcomes and produces 63\% fewer false positives. We further demonstrate that model errors arise from distinct linguistic mechanisms, with false positives driven by discourse-level ambiguity rather than explicit bias cues. These findings highlight the importance of interpretability-aware evaluation for bias detection systems and suggest that architectural and training choices critically affect both model reliability and deployment suitability in journalistic contexts.

偏见检测可解释性TransformerSHAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。