用RoBERTa模型检测新闻句子级偏见,提升准确率与可解释性。
To Bias or Not to Bias: Detecting bias in News with bias-detector
- 基于专家标注的BABE数据集微调RoBERTa模型进行句级偏见分类。
- 相比基线模型性能显著提升,统计检验结果具显著性。
- 注意力分析显示模型更关注上下文而非敏感词汇,避免误判。
媒体偏见检测对保障信息公平传播至关重要,但受限于偏见的主观性及高质量标注数据稀缺。本文通过在专家标注的BABE数据集上微调基于RoBERTa的模型,实现句子级偏见分类。使用McNemar检验和5x2交叉验证配对t检验,证明本模型相比领域自适应预训练的DA-RoBERTa基线有显著性能提升。注意力分析表明,模型未过度敏感于政治敏感词,而是更关注上下文相关词元。为全面评估偏见,提出结合现有偏见类型分类器的流水线方法。尽管受限于句级分析和数据集规模,该方法仍展现出良好泛化性和可解释性。未来方向包括上下文感知建模、偏见中和及高级偏见类型分类。研究成果有助于构建更鲁棒、可解释且具有社会责任感的自然语言处理系统。
原文摘要 · Abstract (English)
Media bias detection is a critical task in ensuring fair and balanced information dissemination, yet it remains challenging due to the subjectivity of bias and the scarcity of high-quality annotated data. In this work, we perform sentence-level bias classification by fine-tuning a RoBERTa-based model on the expert-annotated BABE dataset. Using McNemar's test and the 5x2 cross-validation paired t-test, we show statistically significant improvements in performance when comparing our model to a domain-adaptively pre-trained DA-RoBERTa baseline. Furthermore, attention-based analysis shows that our model avoids common pitfalls like oversensitivity to politically charged terms and instead attends more meaningfully to contextually relevant tokens. For a comprehensive examination of media bias, we present a pipeline that combines our model with an already-existing bias-type classifier. Our method exhibits good generalization and interpretability, despite being constrained by sentence-level analysis and dataset size because of a lack of larger and more advanced bias corpora. We talk about context-aware modeling, bias neutralization, and advanced bias type classification as potential future directions. Our findings contribute to building more robust, explainable, and socially responsible NLP systems for media bias detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。