结合主题模型与大模型,自动识别新闻叙事随时间的演变节点。
Narrative Shift Detection: A Hybrid Approach of Dynamic Topic Models and Large Language Models
- 用主题模型定位关键话题变化点,再筛选相关文本
- 大模型自动解析变化内容,准确识别叙事转变
- 适合关注舆论演化、政策话语分析的研究者
随着媒体叙事快速演变,仅提取文本中的叙事已不够,更需追踪其随时间的发展。现有基于大语言模型的叙事提取方法虽能捕捉典型叙事元素或复杂结构,但在全量语料上应用时存在高昂的计算或财务成本。本文提出一种混合方法:结合大语言模型的语言理解能力与主题模型的大规模适用性,基于叙事政策框架动态建模叙事演变。先用主题模型与变点检测方法识别特定话题的关键变化点,再筛选出代表性文档输入大语言模型,由其自动解释变化并区分内容变化与叙事转变。该方法应用于2009至2023年《华尔街日报》文章语料,结果表明:当某时刻确实存在叙事转变时,大语言模型可高效识别;但在判断是内容变化还是叙事变化方面表现不佳。
原文摘要 · Abstract (English)
With rapidly evolving media narratives, it has become increasingly critical to not just extract narratives from a given corpus but rather investigate, how they develop over time. While popular narrative extraction methods such as Large Language Models do well in capturing typical narrative elements or even the complex structure of a narrative, applying them to an entire corpus comes with obstacles, such as a high financial or computational cost. We propose a combination of the language understanding capabilities of Large Language Models with the large scale applicability of topic models to dynamically model narrative shifts across time using the Narrative Policy Framework. We apply a topic model and a corresponding change point detection method to find changes that concern a specific topic of interest. Using this model, we filter our corpus for documents that are particularly representative of that change and feed them into a Large Language Model that interprets the change that happened in an automated fashion and distinguishes between content and narrative shifts. We employ our pipeline on a corpus of The Wall Street Journal news paper articles from 2009 to 2023. Our findings indicate that a Large Language Model can efficiently extract a narrative shift if one exists at a given point in time, but does not perform as well when having to decide whether a shift in content or a narrative shift took place.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。