追踪4000+网站18个月新闻叙事传播路径,识别虚假信息源头与影响力节点。
Tracking the Takes and Trajectories of English-Language News Narratives across Trustworthy and Worrisome Websites
- 用大模型零样本立场检测,自动识别跨网站新闻叙事与态度。
- 追踪14.6万条新闻,发现反疫苗、反乌等偏颇传播网络。
- 适合舆情分析、事实核查及平台治理研究者使用。
理解误导性与虚假信息如何渗入新闻生态仍具挑战,需追踪其在数千个边缘与主流新闻网站间的传播。本文提出一个系统,利用基于编码器的大语言模型和零样本立场检测,可规模化识别并追踪超过4,000家英文新闻网站(包括事实不可靠、混合可靠性与事实可靠)上的新闻叙事及其立场。系统运行18个月,追踪了14.6万条新闻故事。通过基于网络的干扰分析(NETINF算法),我们发现新闻叙事传播路径与网站对特定实体的立场,可用于揭示偏颇宣传网络(如反疫苗、反乌克兰)并识别在更广泛新闻生态系统中最具影响力的传播网站。提升对分布式新闻生态的可见性,有助于推动宣传与虚假信息的报告与核查。
原文摘要 · Abstract (English)
Understanding how misleading and outright false information enters news ecosystems remains a difficult challenge that requires tracking how narratives spread across thousands of fringe and mainstream news websites. To do this, we introduce a system that utilizes encoder-based large language models and zero-shot stance detection to scalably identify and track news narratives and their attitudes across over 4,000 factually unreliable, mixed-reliability, and factually reliable English-language news websites. Running our system over an 18 month period, we track the spread of 146K news stories. Using network-based interference via the NETINF algorithm, we show that the paths of news narratives and the stances of websites toward particular entities can be used to uncover slanted propaganda networks (e.g., anti-vaccine and anti-Ukraine) and to identify the most influential websites in spreading these attitudes in the broader news ecosystem. We hope that increased visibility into our distributed news ecosystem can help with the reporting and fact-checking of propaganda and disinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。