arXiv:2509.25649cs.CL2025-09被引 1

构建可扩展的新闻偏见分析框架,实时追踪报道选择与话语框架

The Media Bias Detector: A Framework for Annotating and Analyzing the News at Scale

  • 用大模型+爬虫实现每日数百篇新闻的结构化标注
  • 量化句子、文章、媒体三层面的政治倾向与语气特征
  • 适合研究媒体偏见、信息生态的学者与公众监督者

主流新闻机构不仅通过发布文章影响公众认知,还通过选择报道话题(或忽略)以及议题表述方式间接塑造舆论。然而,大规模测量这些细微的媒体偏见仍具挑战。本文提出一个自2024年1月1日起持续运行的近实时数据集与计算框架,支持对新闻报道中的选择性与框架性偏见进行系统研究。该流程结合大语言模型(LLMs)与可扩展的近实时新闻抓取技术,每日自动提取数百篇文章的结构化标注信息,包括政治倾向、语调、主题、文章类型及重大事件等。我们从句级、文章级到媒体级多维度量化新闻覆盖特征,拓展了现代新闻生态中媒体偏见分析的路径。除提供经筛选的数据集外,还开放交互式网页平台,便于数据探索。这些成果建立了一套可复用的大规模媒体偏见研究方法,为未来研究提供实证资源。基于2024年已分析的15万+篇文章,我们展示该数据集在揭示新闻报道模式与偏见方面的潜力,支持学术研究与提升媒体问责的实践。

原文摘要 · Abstract (English)

Mainstream news organizations shape public perception not only directly through the articles they publish but also through the choices they make about which topics to cover (or ignore) and how to frame the issues they do decide to cover. However, measuring these subtle forms of media bias at scale remains a challenge. Here, we introduce a large, ongoing (from January 1, 2024 to present), near real-time dataset and computational framework developed to enable systematic study of selection and framing bias in news coverage. Our pipeline integrates large language models (LLMs) with scalable, near-real-time news scraping to extract structured annotations -- including political lean, tone, topics, article type, and major events -- across hundreds of articles per day. We quantify these dimensions of coverage at multiple levels -- the sentence level, the article level, and the publisher level -- expanding the ways in which researchers can analyze media bias in the modern news landscape. In addition to a curated dataset, we also release an interactive web platform for convenient exploration of these data. Together, these contributions establish a reusable methodology for studying media bias at scale, providing empirical resources for future research. Leveraging the breadth of the corpus over time and across publishers, we also present some examples (focused on the 150,000+ articles examined in 2024) that illustrate how this novel data set can reveal insightful patterns in news coverage and bias, supporting academic research and real-world efforts to improve media accountability.

媒体偏见新闻分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。