arXiv:2412.02271cs.CL2024-12AAAI被引 1

构建新闻标题修改数据集,揭示媒体偏见的动态演化与传播规律

The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias

  • 基于大模型与人工校验,标注7.8万组标题修订及13类偏见
  • 发现编辑后偏见内容在社交平台传播更广,区域框架差异显著
  • 适合研究媒体偏见、信息传播与算法影响的学者和从业者

我们提出MediaSpin数据集,记录主流媒体在发布后对新闻标题的修改情况,并构建了配套的MediaSpin-in-the-Wild数据集,将这些修订标题与社交媒体上的用户互动行为关联。该数据集包含78,910对标题,按13种媒体偏见类型进行标注,涵盖主观(如煽动性、误导性)和客观(如遗漏、倾斜)形式,采用人机协同的大语言模型标注流程并经专家验证。研究展示了三项下游应用:(1) 跨国分析标题中国家引用的增删模式;(2) 使用Transformer模型实现二分类与细粒度偏见识别;(3) 基于819名用户发布的180,786条新闻相关推文,分析偏见标题在X(Twitter)平台的行为影响。结果显示,不同地区存在代表性框架差异,偏见内容具有可测量的语言特征,且持续获得更高互动率。MediaSpin与MediaSpin-in-the-Wild共同构成可复现的偏见检测与媒体生态动态研究基准。

原文摘要 · Abstract (English)

We present MediaSpin, a large-scale language resource capturing how major news outlets modify headlines after publication, and MediaSpin-in-the-Wild, a complementary dataset linking these revised headlines to their downstream engagement on social media. The increasing editability of online news headlines offers new opportunities to study linguistic framing and bias through the lens of editorial revisions. The dataset contains 78,910 headline pairs annotated for 13 types of media bias, grounded in established media-bias taxonomies, covering both subjective (e.g., sensationalism, spin) and objective (e.g., omission, slant) forms, with annotation conducted through a human-supervised large-language-model pipeline with expert validation and quality control. We describe the annotation schema and demonstrate three downstream applications: (1) cross-national analysis of how country references are added or removed during editing, (2) transformer-based bias classification at both binary and fine-grained levels, and (3) behavioral analysis of biased headlines on X (Twitter) using 180,786 news-related tweets from 819 consenting users. The results reveal regional asymmetries in representational framing, measurable linguistic markers, and consistently higher engagement with biased content. MediaSpin and MediaSpin-in-the-Wild together provide a reproducible benchmark for bias detection and the study of editorial and behavioral dynamics in contemporary media ecosystems.

媒体偏见数据集自然语言处理社交传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。