arXiv:2608.17454cs.CL2026-08

用大模型分析新闻偏见,自动识别人物提及与语气差异。

From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis

  • 构建端到端管道,提取实体与情感,比较不同媒体表述
  • 在8358篇阿尔巴尼亚新闻上验证,情感与实体识别有中等一致性
  • 适合缺乏人工标注数据的地区或语言场景使用

本文提出一种分析在线新闻媒体偏见与框架的流程。该流程将文章按主题和事件分组,添加命名实体与情感标注,并通过人物提及、媒体层级语气及事件覆盖模式对比不同新闻源。研究基于从GDELT收集的8,358篇阿尔巴尼亚新闻,与GDELT的自动标注结果进行比对。结果显示,情感与实体抽取达到中等一致度,并发现额外的人物-实体关联对偏见分析具有潜在价值。比较两种标注提示后发现,更严格的验证规则虽减少标签不一致,但延长执行时间并降低覆盖率;因此后续采用较简单提示。样本分析展示了媒体层级框架特征、人物语气差异及事件覆盖集中度等指标,表明同一数据集可用于揭示报道内容、描述方式与覆盖焦点。该方法在缺乏人工标注数据或专用语言工具时尤为适用。

原文摘要 · Abstract (English)

This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level coverage patterns. We apply it to 8,358 Albanian news articles collected from GDELT and compare the resulting annotations with GDELT's automated annotations. The results show moderate agreement for sentiment and entity extraction, as well as additional person-entity pairs that can potentially support the bias analysis. We compare two annotation prompts and find that stricter sentiment-validation rules remove label-score inconsistencies but increase execution time and reduce annotation coverage. Based on these results, the simpler prompt is used for the rest of the analysis. We have provided sample analysis on source-level framing pro les, person-level tone differences across sources, and event-level gatekeeping and coverage indicators. These outputs show how the same news collection can be used to examine what sources cover, how they describe public figures, and where coverage is concentrated. The approach is particularly useful in settings where manually verified datasets or specialized language tools are limited.

媒体偏见大模型文本分析新闻语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。