评估新闻偏见检测模型解释力的多维指标,发现模型表现、解释可信度和机制忠实性需分开考量。
A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

- 用专家标注的BABE数据集,从预测性能、解释可信度、机制忠实性三方面评估模型
- 注意力监督微调提升解释可信度,但不同模型架构表现差异大
- 模型规模不决定可压缩性,机制分析揭示架构间显著差异
自动检测媒体偏见困难,因偏见常隐含于细微语境中,仅准确预测不足,还需反映模型推理过程的解释。本文基于专家标注的BABE数据集,对BERT与RoBERTa(base与large版本)在编码器型媒体偏见检测中的可解释性进行多维度评估。从三个互补维度展开:预测性能、解释合理性(标记层面与专家理由的对齐度)、机制忠实性(在反事实理由遮蔽下,紧凑注意力头集合能否恢复预测信号)。为引入解释合理性变化,额外研究注意力监督微调,即以专家理由作为辅助训练信号。注意力监督作为对归因合理性的干预手段,而归因方法效果在不同架构间差异显著。电路分析进一步显示,机制可恢复性在不同架构间存在巨大差异,表明模型规模本身并不决定电路可压缩性。综合结果表明,预测性能、归因合理性与机制忠实性分别刻画了模型行为的不同侧面,应独立评估以深入理解媒体偏见检测中的可解释性。
原文摘要 · Abstract (English)
Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone are insufficient without explanations that reflect the model's underlying reasoning. We present a multi-dimensional evaluation of explainability in encoder-based media bias detection using the Bias Annotations By Experts (BABE) dataset. Specifically, we study BERT and RoBERTa as classifiers (base and large variants) along three complementary axes: predictive performance, explanation plausibility (token-level alignment with expert rationales), and mechanistic faithfulness (whether compact sets of attention heads recover predictive signal under counterfactual rationale masking). To induce variation in plausibility, we additionally investigate attention-supervised finetuning, which incorporates expert rationale annotations as an auxiliary training signal. Attention supervision serves as an intervention on attribution plausibility, while the effectiveness of attribution methods varies substantially across architectures. Circuit analysis further reveals substantial variation in mechanistic recoverability across architectures, suggesting that model scale alone does not determine circuit compressibility. Taken together, our findings suggest that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。