arXiv:2410.17655cs.AIcs.CY2024-10中稿 · CLEF 2024被引 3

通过网页互动分析新闻媒体的客观性和政治偏见,效果优于现有方法。

Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions

  • 基于新闻媒体间的超链接关系建模,用强化学习识别偏见。
  • 在事实性与政治偏见预测上均显著提升性能,F1和MAE更优。
  • 开源最大规模标注数据集,适合媒体研究与信息可信度评估者使用。

新闻来源的偏见评估对依赖真实信息的专业人士、机构和研究人员至关重要。尽管部分内容特征可识别偏见,但‘事实性报道’与‘政治偏见’等描述仍具挑战。本文扩展了一种近期提出的新闻媒体可靠性评估方法,聚焦于媒体及其长期网络互动关系的建模。具体而言,我们在大规模新闻媒体超链接图上评估了四种强化学习策略的分类性能。实验针对‘事实性报道’和‘政治偏见’两个难题,结果表明在媒体源层面有显著性能提升。此外,我们在 CLEF 2023 CheckThat! Lab 挑战中验证方法,其 F1 分数和官方 MAE 指标均优于已有报告结果。我们还贡献了目前最大的标注新闻媒体数据集,包含事实性与政治偏见标签。研究结果表明,基于时间演化的超链接互动对新闻媒体进行画像是可行的,为理解动态媒体格局提供了全景视角。

原文摘要 · Abstract (English)

Bias assessment of news sources is paramount for professionals, organizations, and researchers who rely on truthful evidence for information gathering and reporting. While certain bias indicators are discernible from content analysis, descriptors like political bias and fake news pose greater challenges. In this paper, we propose an extension to a recently presented news media reliability estimation method that focuses on modeling outlets and their longitudinal web interactions. Concretely, we assess the classification performance of four reinforcement learning strategies on a large news media hyperlink graph. Our experiments, targeting two challenging bias descriptors, factual reporting and political bias, showed a significant performance improvement at the source media level. Additionally, we validate our methods on the CLEF 2023 CheckThat! Lab challenge, outperforming the reported results in both, F1-score and the official MAE metric. Furthermore, we contribute by releasing the largest annotated dataset of news source media, categorized with factual reporting and political bias labels. Our findings suggest that profiling news media sources based on their hyperlink interactions over time is feasible, offering a bird's-eye view of evolving media landscapes.

媒体偏见网络分析强化学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。